Skip to content

Model Pricing Comparison

Auto-generated on 2026-05-31 from OpenRouter API (/api/v1/models). All prices in USD per 1M tokens. β€œOut/In” column shows the output-to-input price ratio β€” the hidden cost multiplier.


Your current default model is owl-alpha β€” completely free.

ModelInputOutputCtxProvider
owl-alphaFREEFREE1Mopenrouter

🟒 Ultra-Cheap Models (<$0.15 input, good out/in ratio)

Section titled β€œπŸŸ’ Ultra-Cheap Models (<$0.15 input, good out/in ratio)”

These are the everyday workhorses β€” cheap enough for casual use without watching the meter.

ModelInput/MOutput/MOut/InCtxProvider
llama-3.1-8b-instruct$0.020$0.0502.5x131Kmeta-llama
llama-3-8b-instruct$0.040$0.0401.0x8Kmeta-llama
qwen-2.5-7b-instruct$0.040$0.1002.5x131Kqwen
qwen3.5-9b$0.040$0.1503.8x262Kqwen
gemma-3-4b-it$0.040$0.0802.0x131Kgoogle
gemma-3-12b-it$0.040$0.1303.2x131Kgoogle
gemma-3-27b-it$0.080$0.1602.0x131Kgoogle
gemma-3n-e4b-it$0.060$0.1202.0x32Kgoogle
mistral-nemo$0.020$0.0301.5x131Kmistralai
mistral-small-24b-instruct-2501$0.050$0.0801.6x32Kmistralai
mistral-small-3.2-24b-instruct$0.075$0.2002.7x128Kmistralai
gpt-oss-20b$0.029$0.1404.8x131Kopenai
gpt-oss-120b$0.039$0.1804.6x131Kopenai
gpt-4.1-nano$0.100$0.4004.0x1Mopenai
gpt-5-nano$0.050$0.4008.0x400Kopenai
deepseek-v4-flash$0.098$0.1972.0x1Mdeepseek
deepseek-r1-distill-llama-70b$0.700$0.8001.1x131Kdeepseek
deepseek-r1-distill-qwen-32b$0.290$0.2901.0x128Kdeepseek
qwen3-235b-a22b-2507$0.071$0.1001.4x262Kqwen
qwen3-coder-30b-a3b-instruct$0.070$0.2703.9x160Kqwen
qwen3.5-flash-02-23$0.065$0.2604.0x1Mqwen
qwen3-14b$0.100$0.2402.4x131Kqwen
qwen3-32b$0.080$0.2803.5x131Kqwen
qwen3-8b$0.050$0.4008.0x131Kqwen
llama-4-scout$0.080$0.3003.8x10Mmeta-llama
ministral-3b-2512$0.100$0.1001.0x131Kmistralai
ministral-8b-2512$0.150$0.1501.0x262Kmistralai
devstral-small$0.100$0.3003.0x131Kmistralai
gemini-2.0-flash-lite-001$0.075$0.3004.0x1Mgoogle
gemini-2.0-flash-001$0.100$0.4004.0x1Mgoogle
gemini-2.5-flash-lite$0.100$0.4004.0x1Mgoogle
gemini-2.5-flash-lite-preview-09-2025$0.100$0.4004.0x1Mgoogle
step-3.5-flash$0.090$0.3003.3x262Kstepfun
gpt-4o-mini$0.150$0.6004.0x128Kopenai
gpt-4.1-mini$0.400$1.6004.0x1Mopenai

Still reasonable for production use, but output costs start adding up on long conversations.

ModelInput/MOutput/MOut/InCtxProvider
deepseek-v3.2$0.252$0.3781.5x131Kdeepseek
deepseek-v3.2-exp$0.270$0.4101.5x163Kdeepseek
deepseek-chat-v3-0324$0.200$0.7703.9x163Kdeepseek
deepseek-chat-v3.1$0.210$0.7903.8x163Kdeepseek
deepseek-chat$0.229$0.9144.0x131Kdeepseek
deepseek-v3.1-terminus$0.270$0.9503.5x163Kdeepseek
deepseek-v4-pro$0.435$0.8702.0x1Mdeepseek
qwen-plus$0.260$0.7803.0x1Mqwen
qwen-plus-2025-07-28$0.260$0.7803.0x1Mqwen
qwen3.5-plus-02-15$0.260$1.5606.0x1Mqwen
qwen3.5-plus-20260420$0.300$1.8006.0x1Mqwen
qwen3.6-plus$0.325$1.9506.0x1Mqwen
qwen3.6-35b-a3b$0.140$1.0007.1x262Kqwen
qwen3.5-35b-a3b$0.140$1.0007.1x262Kqwen
qwen3.5-27b$0.195$1.5608.0x262Kqwen
qwen3-235b-a22b$0.455$1.8204.0x131Kqwen
claude-3-haiku$0.250$1.2505.0x200Kanthropic
gpt-5.4-nano$0.200$1.2506.3x400Kopenai
gpt-5-mini$0.250$2.0008.0x400Kopenai
gpt-5.4-mini$0.750$4.5006.0x400Kopenai
step-3.7-flash$0.200$1.1505.8x256Kstepfun
gemini-2.5-flash$0.300$2.5008.3x1Mgoogle
gemini-2.5-flash-image$0.300$2.5008.3x32Kgoogle
gemini-3.1-flash-lite$0.250$1.5006.0x1Mgoogle
gemini-3.1-flash-lite-preview$0.250$1.5006.0x1Mgoogle
gpt-3.5-turbo$0.500$1.5003.0x16Kopenai
gpt-3.5-turbo-0613$1.000$2.0002.0x4Kopenai
gpt-3.5-turbo-16k$3.000$4.0001.3x16Kopenai
gpt-3.5-turbo-instruct$1.500$2.0001.3x4Kopenai

πŸ”΄ The Trap β€” Reasoning Models & High Output/Input Ratio

Section titled β€œπŸ”΄ The Trap β€” Reasoning Models & High Output/Input Ratio”

These models look cheap on input but the output costs 3.5–4.3x more. One long reasoning trace burns through tokens fast. This is the most likely cause of your surprise €2–3 DeepSeek bills.

ModelInput/MOutput/MOut/InCtxProvider
deepseek-r1-0528$0.500$2.1504.3x163Kdeepseek
deepseek-r1$0.700$2.5003.6x163Kdeepseek
qwen3-30b-a3b-thinking-2507$0.080$0.4005.0x131Kqwen
qwen3-next-80b-a3b-thinking$0.098$0.7808.0x262Kqwen
qwen3-235b-a22b-thinking-2507$0.150$1.50010.0x262Kqwen
qwen3-vl-8b-thinking$0.117$1.36511.7x256Kqwen
qwen3-vl-30b-a3b-thinking$0.130$1.56012.0x131Kqwen
qwen3-vl-235b-a22b-thinking$0.260$2.60010.0x131Kqwen
qwen3-max-thinking$0.780$3.9005.0x262Kqwen
o3-mini$1.100$4.4004.0x200Kopenai
o4-mini$1.100$4.4004.0x200Kopenai
o3$2.000$8.0004.0x200Kopenai
o1$15.000$60.0004.0x200Kopenai

Consistent 5x output-to-input ratio across the board. Clean pricing, no surprises β€” but premium absolute prices.

ModelInput/MOutput/MOut/InCtx
claude-3-haiku$0.250$1.2505.0x200K
claude-3.5-haiku$0.800$4.0005.0x200K
claude-haiku-4.5$1.000$5.0005.0x200K
claude-sonnet-4.6$3.000$15.0005.0x1M
claude-sonnet-4.5$3.000$15.0005.0x1M
claude-sonnet-4$3.000$15.0005.0x1M
claude-opus-4.8$5.000$25.0005.0x1M
claude-opus-4.7$5.000$25.0005.0x1M
claude-opus-4.6$5.000$25.0005.0x1M
claude-opus-4$15.000$75.0005.0x200K
claude-opus-4.1$15.000$75.0005.0x200K
claude-opus-4.8-fast ⚑$10.000$50.0005.0x1M
claude-opus-4.7-fast ⚑$30.000$150.0005.0x1M
claude-opus-4.6-fast ⚑$30.000$150.0005.0x1M

⚑ = fast mode variant, 2Γ— the regular price for higher throughput


ModelInput/MOutput/MOut/InCtx
gpt-4.1-nano$0.100$0.4004.0x1M
gpt-4o-mini$0.150$0.6004.0x128K
gpt-4.1-mini$0.400$1.6004.0x1M
gpt-4.1$2.000$8.0004.0x1M
gpt-4o$2.500$10.0004.0x128K
gpt-4o-2024-08-06$2.500$10.0004.0x128K
gpt-4o-search-preview$2.500$10.0004.0x128K
gpt-5-nano$0.050$0.4008.0x400K
gpt-5-mini$0.250$2.0008.0x400K
gpt-5$1.250$10.0008.0x400K
gpt-5-chat$1.250$10.0008.0x128K
gpt-5.1$1.250$10.0008.0x400K
gpt-5.4$2.500$15.0006.0x1M
gpt-5.5$5.000$30.0006.0x1M
gpt-5-pro$15.000$120.0008.0x400K
gpt-5.2$1.750$14.0008.0x400K
gpt-5.3$1.750$14.0008.0x400K
gpt-5.2-pro$21.000$168.0008.0x400K
gpt-5.4-pro$30.000$180.0006.0x1M
gpt-5.5-pro$30.000$180.0006.0x1M
gpt-4-turbo$10.000$30.0003.0x128K
gpt-4$30.000$60.0002.0x8K
gpt-chat-latest$5.000$30.0006.0x400K

ModelInput/MOutput/MOut/InCtx
gemini-2.0-flash-lite-001$0.075$0.3004.0x1M
gemini-2.0-flash-001$0.100$0.4004.0x1M
gemini-2.5-flash-lite$0.100$0.4004.0x1M
gemini-2.5-flash$0.300$2.5008.3x1M
gemini-3.1-flash-lite$0.250$1.5006.0x1M
gemini-2.5-pro$1.250$10.0008.0x1M
gemini-3.1-pro-preview$2.000$12.0006.0x1M
gemini-3.5-flash$1.500$9.0006.0x1M
gemini-3-flash-preview$0.500$3.0006.0x1M
gemini-3.1-flash-image-preview$0.500$3.0006.0x131K

Assumptions: ~5K context tokens per turn (full conversation history), ~1K output tokens per turn. Total: ~50K input + ~10K output.

This is the realistic Hermes session pattern β€” each turn sends the full context window.

ModelInput CostOutput CostTotal
owl-alphaFREEFREE€0.00
deepseek-v4-flash$0.005$0.002€0.007
qwen3-235b-a22b-2507$0.004$0.001€0.005
gpt-4.1-nano$0.005$0.004€0.009
deepseek-v3.2$0.013$0.004€0.016
qwen3.6-35b-a3b$0.007$0.010€0.017
deepseek-chat-v3-0324$0.010$0.008€0.018
gpt-4o-mini$0.008$0.006€0.014
gpt-5-mini$0.013$0.020€0.033
deepseek-r1-0528$0.025$0.022€0.047
claude-3-haiku$0.013$0.013€0.025
deepseek-r1$0.035$0.025€0.060
claude-haiku-4.5$0.050$0.050€0.10
gpt-4.1$0.100$0.080€0.18
gpt-4o$0.125$0.100€0.23
o3-mini$0.055$0.044€0.10
claude-sonnet-4.5$0.150$0.150€0.30
gpt-5$0.063$0.100€0.16
claude-opus-4.8$0.250$0.250€0.50
  • The out/in ratio is the hidden cost. Models like o3-mini and deepseek-r1 have 4x+ output multipliers. A 2K-token response at $4.40/M costs $0.009 β€” multiply that by 10+ turns with long reasoning traces and you hit €2–3 fast.
  • Context window = compounding cost. Hermes sends full conversation history each turn. At 10K context Γ— 20 turns = 200K input tokens. R1 input at $0.70/M = $0.14 for input alone before any output.
  • Free models are legitimately free. owl-alpha at $0/M means your Hermes sessions cost nothing in LLM fees.
  • Claude pricing is predictable (always 5x output) but premium. Sonnet 4.5 for a heavy session runs ~€0.30, Opus ~€0.50.
  • The €2–3 mystery solved: Almost certainly deepseek-r1 or deepseek-chat-v3-0324 with verbose output, compounding over multiple turns with full context. Not a bug β€” just the hidden cost of reasoning models.