Skip to content

Model Pricing Comparison

Auto-generated on 2026-05-31 from OpenRouter API (/api/v1/models). All prices in USD per 1M tokens. β€œOut/In” column shows the output-to-input price ratio β€” the hidden cost multiplier.


Model Input Output Ctx Provider
owl-alpha FREE FREE 1M openrouter

🟒 Ultra-Cheap Models (<$0.15 input, good out/in ratio)

Section titled β€œπŸŸ’ Ultra-Cheap Models (<$0.15 input, good out/in ratio)”

These are the everyday workhorses β€” cheap enough for casual use without watching the meter.

Model Input/M Output/M Out/In Ctx Provider
llama-3.1-8b-instruct $0.020 $0.050 2.5x 131K meta-llama
llama-3-8b-instruct $0.040 $0.040 1.0x 8K meta-llama
qwen-2.5-7b-instruct $0.040 $0.100 2.5x 131K qwen
qwen3.5-9b $0.040 $0.150 3.8x 262K qwen
gemma-3-4b-it $0.040 $0.080 2.0x 131K google
gemma-3-12b-it $0.040 $0.130 3.2x 131K google
gemma-3-27b-it $0.080 $0.160 2.0x 131K google
gemma-3n-e4b-it $0.060 $0.120 2.0x 32K google
mistral-nemo $0.020 $0.030 1.5x 131K mistralai
mistral-small-24b-instruct-2501 $0.050 $0.080 1.6x 32K mistralai
mistral-small-3.2-24b-instruct $0.075 $0.200 2.7x 128K mistralai
gpt-oss-20b $0.029 $0.140 4.8x 131K openai
gpt-oss-120b $0.039 $0.180 4.6x 131K openai
gpt-4.1-nano $0.100 $0.400 4.0x 1M openai
gpt-5-nano $0.050 $0.400 8.0x 400K openai
deepseek-v4-flash $0.098 $0.197 2.0x 1M deepseek
deepseek-r1-distill-llama-70b $0.700 $0.800 1.1x 131K deepseek
deepseek-r1-distill-qwen-32b $0.290 $0.290 1.0x 128K deepseek
qwen3-235b-a22b-2507 $0.071 $0.100 1.4x 262K qwen
qwen3-coder-30b-a3b-instruct $0.070 $0.270 3.9x 160K qwen
qwen3.5-flash-02-23 $0.065 $0.260 4.0x 1M qwen
qwen3-14b $0.100 $0.240 2.4x 131K qwen
qwen3-32b $0.080 $0.280 3.5x 131K qwen
qwen3-8b $0.050 $0.400 8.0x 131K qwen
llama-4-scout $0.080 $0.300 3.8x 10M meta-llama
ministral-3b-2512 $0.100 $0.100 1.0x 131K mistralai
ministral-8b-2512 $0.150 $0.150 1.0x 262K mistralai
devstral-small $0.100 $0.300 3.0x 131K mistralai
gemini-2.0-flash-lite-001 $0.075 $0.300 4.0x 1M google
gemini-2.0-flash-001 $0.100 $0.400 4.0x 1M google
gemini-2.5-flash-lite $0.100 $0.400 4.0x 1M google
gemini-2.5-flash-lite-preview-09-2025 $0.100 $0.400 4.0x 1M google
step-3.5-flash $0.090 $0.300 3.3x 262K stepfun
gpt-4o-mini $0.150 $0.600 4.0x 128K openai
gpt-4.1-mini $0.400 $1.600 4.0x 1M openai

Still reasonable for production use, but output costs start adding up on long conversations.

Model Input/M Output/M Out/In Ctx Provider
deepseek-v3.2 $0.252 $0.378 1.5x 131K deepseek
deepseek-v3.2-exp $0.270 $0.410 1.5x 163K deepseek
deepseek-chat-v3-0324 $0.200 $0.770 3.9x 163K deepseek
deepseek-chat-v3.1 $0.210 $0.790 3.8x 163K deepseek
deepseek-chat $0.229 $0.914 4.0x 131K deepseek
deepseek-v3.1-terminus $0.270 $0.950 3.5x 163K deepseek
deepseek-v4-pro $0.435 $0.870 2.0x 1M deepseek
qwen-plus $0.260 $0.780 3.0x 1M qwen
qwen-plus-2025-07-28 $0.260 $0.780 3.0x 1M qwen
qwen3.5-plus-02-15 $0.260 $1.560 6.0x 1M qwen
qwen3.5-plus-20260420 $0.300 $1.800 6.0x 1M qwen
qwen3.6-plus $0.325 $1.950 6.0x 1M qwen
qwen3.6-35b-a3b $0.140 $1.000 7.1x 262K qwen
qwen3.5-35b-a3b $0.140 $1.000 7.1x 262K qwen
qwen3.5-27b $0.195 $1.560 8.0x 262K qwen
qwen3-235b-a22b $0.455 $1.820 4.0x 131K qwen
claude-3-haiku $0.250 $1.250 5.0x 200K anthropic
gpt-5.4-nano $0.200 $1.250 6.3x 400K openai
gpt-5-mini $0.250 $2.000 8.0x 400K openai
gpt-5.4-mini $0.750 $4.500 6.0x 400K openai
step-3.7-flash $0.200 $1.150 5.8x 256K stepfun
gemini-2.5-flash $0.300 $2.500 8.3x 1M google
gemini-2.5-flash-image $0.300 $2.500 8.3x 32K google
gemini-3.1-flash-lite $0.250 $1.500 6.0x 1M google
gemini-3.1-flash-lite-preview $0.250 $1.500 6.0x 1M google
gpt-3.5-turbo $0.500 $1.500 3.0x 16K openai
gpt-3.5-turbo-0613 $1.000 $2.000 2.0x 4K openai
gpt-3.5-turbo-16k $3.000 $4.000 1.3x 16K openai
gpt-3.5-turbo-instruct $1.500 $2.000 1.3x 4K openai

πŸ”΄ The Trap β€” Reasoning Models & High Output/Input Ratio

Section titled β€œπŸ”΄ The Trap β€” Reasoning Models & High Output/Input Ratio”

These models look cheap on input but the output costs 3.5–4.3x more. One long reasoning trace burns through tokens fast. This is the most likely cause of your surprise €2–3 DeepSeek bills.

Model Input/M Output/M Out/In Ctx Provider
deepseek-r1-0528 $0.500 $2.150 4.3x 163K deepseek
deepseek-r1 $0.700 $2.500 3.6x 163K deepseek
qwen3-30b-a3b-thinking-2507 $0.080 $0.400 5.0x 131K qwen
qwen3-next-80b-a3b-thinking $0.098 $0.780 8.0x 262K qwen
qwen3-235b-a22b-thinking-2507 $0.150 $1.500 10.0x 262K qwen
qwen3-vl-8b-thinking $0.117 $1.365 11.7x 256K qwen
qwen3-vl-30b-a3b-thinking $0.130 $1.560 12.0x 131K qwen
qwen3-vl-235b-a22b-thinking $0.260 $2.600 10.0x 131K qwen
qwen3-max-thinking $0.780 $3.900 5.0x 262K qwen
o3-mini $1.100 $4.400 4.0x 200K openai
o4-mini $1.100 $4.400 4.0x 200K openai
o3 $2.000 $8.000 4.0x 200K openai
o1 $15.000 $60.000 4.0x 200K openai

Consistent 5x output-to-input ratio across the board. Clean pricing, no surprises β€” but premium absolute prices.

Model Input/M Output/M Out/In Ctx
claude-3-haiku $0.250 $1.250 5.0x 200K
claude-3.5-haiku $0.800 $4.000 5.0x 200K
claude-haiku-4.5 $1.000 $5.000 5.0x 200K
claude-sonnet-4.6 $3.000 $15.000 5.0x 1M
claude-sonnet-4.5 $3.000 $15.000 5.0x 1M
claude-sonnet-4 $3.000 $15.000 5.0x 1M
claude-opus-4.8 $5.000 $25.000 5.0x 1M
claude-opus-4.7 $5.000 $25.000 5.0x 1M
claude-opus-4.6 $5.000 $25.000 5.0x 1M
claude-opus-4 $15.000 $75.000 5.0x 200K
claude-opus-4.1 $15.000 $75.000 5.0x 200K
claude-opus-4.8-fast ⚑ $10.000 $50.000 5.0x 1M
claude-opus-4.7-fast ⚑ $30.000 $150.000 5.0x 1M
claude-opus-4.6-fast ⚑ $30.000 $150.000 5.0x 1M

⚑ = fast mode variant, 2Γ— the regular price for higher throughput


Model Input/M Output/M Out/In Ctx
gpt-4.1-nano $0.100 $0.400 4.0x 1M
gpt-4o-mini $0.150 $0.600 4.0x 128K
gpt-4.1-mini $0.400 $1.600 4.0x 1M
gpt-4.1 $2.000 $8.000 4.0x 1M
gpt-4o $2.500 $10.000 4.0x 128K
gpt-4o-2024-08-06 $2.500 $10.000 4.0x 128K
gpt-4o-search-preview $2.500 $10.000 4.0x 128K
gpt-5-nano $0.050 $0.400 8.0x 400K
gpt-5-mini $0.250 $2.000 8.0x 400K
gpt-5 $1.250 $10.000 8.0x 400K
gpt-5-chat $1.250 $10.000 8.0x 128K
gpt-5.1 $1.250 $10.000 8.0x 400K
gpt-5.4 $2.500 $15.000 6.0x 1M
gpt-5.5 $5.000 $30.000 6.0x 1M
gpt-5-pro $15.000 $120.000 8.0x 400K
gpt-5.2 $1.750 $14.000 8.0x 400K
gpt-5.3 $1.750 $14.000 8.0x 400K
gpt-5.2-pro $21.000 $168.000 8.0x 400K
gpt-5.4-pro $30.000 $180.000 6.0x 1M
gpt-5.5-pro $30.000 $180.000 6.0x 1M
gpt-4-turbo $10.000 $30.000 3.0x 128K
gpt-4 $30.000 $60.000 2.0x 8K
gpt-chat-latest $5.000 $30.000 6.0x 400K

Model Input/M Output/M Out/In Ctx
gemini-2.0-flash-lite-001 $0.075 $0.300 4.0x 1M
gemini-2.0-flash-001 $0.100 $0.400 4.0x 1M
gemini-2.5-flash-lite $0.100 $0.400 4.0x 1M
gemini-2.5-flash $0.300 $2.500 8.3x 1M
gemini-3.1-flash-lite $0.250 $1.500 6.0x 1M
gemini-2.5-pro $1.250 $10.000 8.0x 1M
gemini-3.1-pro-preview $2.000 $12.000 6.0x 1M
gemini-3.5-flash $1.500 $9.000 6.0x 1M
gemini-3-flash-preview $0.500 $3.000 6.0x 1M
gemini-3.1-flash-image-preview $0.500 $3.000 6.0x 131K

Assumptions: ~5K context tokens per turn (full conversation history), ~1K output tokens per turn. Total: ~50K input + ~10K output.

This is the realistic Hermes session pattern β€” each turn sends the full context window.

Model Input Cost Output Cost Total
owl-alpha FREE FREE €0.00
deepseek-v4-flash $0.005 $0.002 €0.007
qwen3-235b-a22b-2507 $0.004 $0.001 €0.005
gpt-4.1-nano $0.005 $0.004 €0.009
deepseek-v3.2 $0.013 $0.004 €0.016
qwen3.6-35b-a3b $0.007 $0.010 €0.017
deepseek-chat-v3-0324 $0.010 $0.008 €0.018
gpt-4o-mini $0.008 $0.006 €0.014
gpt-5-mini $0.013 $0.020 €0.033
deepseek-r1-0528 $0.025 $0.022 €0.047
claude-3-haiku $0.013 $0.013 €0.025
deepseek-r1 $0.035 $0.025 €0.060
claude-haiku-4.5 $0.050 $0.050 €0.10
gpt-4.1 $0.100 $0.080 €0.18
gpt-4o $0.125 $0.100 €0.23
o3-mini $0.055 $0.044 €0.10
claude-sonnet-4.5 $0.150 $0.150 €0.30
gpt-5 $0.063 $0.100 €0.16
claude-opus-4.8 $0.250 $0.250 €0.50
  • The out/in ratio is the hidden cost. Models like o3-mini and deepseek-r1 have 4x+ output multipliers. A 2K-token response at $4.40/M costs $0.009 β€” multiplied by 10+ turns with long reasoning traces, it hits €2–3 fast.
  • Context window = compounding cost. Hermes sends full conversation history each turn. At 10K context Γ— 20 turns = 200K input tokens. R1 input at $0.70/M = $0.14 for input alone before any output.
  • Free models are legitimately free. owl-alpha at $0/M means Hermes sessions cost nothing in LLM fees.
  • Claude pricing is predictable (always 5x output) but premium. Sonnet 4.5 for a heavy session runs ~€0.30, Opus ~€0.50.
  • The €2–3 mystery solved: Almost certainly deepseek-r1 or deepseek-chat-v3-0324 with verbose output, compounding over multiple turns with full context. Not a bug β€” just the hidden cost of reasoning models.