Model Pricing Comparison
OpenRouter Model Pricing Comparison
Section titled βOpenRouter Model Pricing ComparisonβAuto-generated on 2026-05-31 from OpenRouter API (
/api/v1/models). All prices in USD per 1M tokens. βOut/Inβ column shows the output-to-input price ratio β the hidden cost multiplier.
π’ Free Models
Section titled βπ’ Free ModelsβYour current default model is owl-alpha β completely free.
| Model | Input | Output | Ctx | Provider |
|---|---|---|---|---|
| owl-alpha | FREE | FREE | 1M | openrouter |
π’ Ultra-Cheap Models (<$0.15 input, good out/in ratio)
Section titled βπ’ Ultra-Cheap Models (<$0.15 input, good out/in ratio)βThese are the everyday workhorses β cheap enough for casual use without watching the meter.
| Model | Input/M | Output/M | Out/In | Ctx | Provider |
|---|---|---|---|---|---|
| llama-3.1-8b-instruct | $0.020 | $0.050 | 2.5x | 131K | meta-llama |
| llama-3-8b-instruct | $0.040 | $0.040 | 1.0x | 8K | meta-llama |
| qwen-2.5-7b-instruct | $0.040 | $0.100 | 2.5x | 131K | qwen |
| qwen3.5-9b | $0.040 | $0.150 | 3.8x | 262K | qwen |
| gemma-3-4b-it | $0.040 | $0.080 | 2.0x | 131K | |
| gemma-3-12b-it | $0.040 | $0.130 | 3.2x | 131K | |
| gemma-3-27b-it | $0.080 | $0.160 | 2.0x | 131K | |
| gemma-3n-e4b-it | $0.060 | $0.120 | 2.0x | 32K | |
| mistral-nemo | $0.020 | $0.030 | 1.5x | 131K | mistralai |
| mistral-small-24b-instruct-2501 | $0.050 | $0.080 | 1.6x | 32K | mistralai |
| mistral-small-3.2-24b-instruct | $0.075 | $0.200 | 2.7x | 128K | mistralai |
| gpt-oss-20b | $0.029 | $0.140 | 4.8x | 131K | openai |
| gpt-oss-120b | $0.039 | $0.180 | 4.6x | 131K | openai |
| gpt-4.1-nano | $0.100 | $0.400 | 4.0x | 1M | openai |
| gpt-5-nano | $0.050 | $0.400 | 8.0x | 400K | openai |
| deepseek-v4-flash | $0.098 | $0.197 | 2.0x | 1M | deepseek |
| deepseek-r1-distill-llama-70b | $0.700 | $0.800 | 1.1x | 131K | deepseek |
| deepseek-r1-distill-qwen-32b | $0.290 | $0.290 | 1.0x | 128K | deepseek |
| qwen3-235b-a22b-2507 | $0.071 | $0.100 | 1.4x | 262K | qwen |
| qwen3-coder-30b-a3b-instruct | $0.070 | $0.270 | 3.9x | 160K | qwen |
| qwen3.5-flash-02-23 | $0.065 | $0.260 | 4.0x | 1M | qwen |
| qwen3-14b | $0.100 | $0.240 | 2.4x | 131K | qwen |
| qwen3-32b | $0.080 | $0.280 | 3.5x | 131K | qwen |
| qwen3-8b | $0.050 | $0.400 | 8.0x | 131K | qwen |
| llama-4-scout | $0.080 | $0.300 | 3.8x | 10M | meta-llama |
| ministral-3b-2512 | $0.100 | $0.100 | 1.0x | 131K | mistralai |
| ministral-8b-2512 | $0.150 | $0.150 | 1.0x | 262K | mistralai |
| devstral-small | $0.100 | $0.300 | 3.0x | 131K | mistralai |
| gemini-2.0-flash-lite-001 | $0.075 | $0.300 | 4.0x | 1M | |
| gemini-2.0-flash-001 | $0.100 | $0.400 | 4.0x | 1M | |
| gemini-2.5-flash-lite | $0.100 | $0.400 | 4.0x | 1M | |
| gemini-2.5-flash-lite-preview-09-2025 | $0.100 | $0.400 | 4.0x | 1M | |
| step-3.5-flash | $0.090 | $0.300 | 3.3x | 262K | stepfun |
| gpt-4o-mini | $0.150 | $0.600 | 4.0x | 128K | openai |
| gpt-4.1-mini | $0.400 | $1.600 | 4.0x | 1M | openai |
π‘ Mid-Tier Models ($0.15β$0.50 input)
Section titled βπ‘ Mid-Tier Models ($0.15β$0.50 input)βStill reasonable for production use, but output costs start adding up on long conversations.
| Model | Input/M | Output/M | Out/In | Ctx | Provider |
|---|---|---|---|---|---|
| deepseek-v3.2 | $0.252 | $0.378 | 1.5x | 131K | deepseek |
| deepseek-v3.2-exp | $0.270 | $0.410 | 1.5x | 163K | deepseek |
| deepseek-chat-v3-0324 | $0.200 | $0.770 | 3.9x | 163K | deepseek |
| deepseek-chat-v3.1 | $0.210 | $0.790 | 3.8x | 163K | deepseek |
| deepseek-chat | $0.229 | $0.914 | 4.0x | 131K | deepseek |
| deepseek-v3.1-terminus | $0.270 | $0.950 | 3.5x | 163K | deepseek |
| deepseek-v4-pro | $0.435 | $0.870 | 2.0x | 1M | deepseek |
| qwen-plus | $0.260 | $0.780 | 3.0x | 1M | qwen |
| qwen-plus-2025-07-28 | $0.260 | $0.780 | 3.0x | 1M | qwen |
| qwen3.5-plus-02-15 | $0.260 | $1.560 | 6.0x | 1M | qwen |
| qwen3.5-plus-20260420 | $0.300 | $1.800 | 6.0x | 1M | qwen |
| qwen3.6-plus | $0.325 | $1.950 | 6.0x | 1M | qwen |
| qwen3.6-35b-a3b | $0.140 | $1.000 | 7.1x | 262K | qwen |
| qwen3.5-35b-a3b | $0.140 | $1.000 | 7.1x | 262K | qwen |
| qwen3.5-27b | $0.195 | $1.560 | 8.0x | 262K | qwen |
| qwen3-235b-a22b | $0.455 | $1.820 | 4.0x | 131K | qwen |
| claude-3-haiku | $0.250 | $1.250 | 5.0x | 200K | anthropic |
| gpt-5.4-nano | $0.200 | $1.250 | 6.3x | 400K | openai |
| gpt-5-mini | $0.250 | $2.000 | 8.0x | 400K | openai |
| gpt-5.4-mini | $0.750 | $4.500 | 6.0x | 400K | openai |
| step-3.7-flash | $0.200 | $1.150 | 5.8x | 256K | stepfun |
| gemini-2.5-flash | $0.300 | $2.500 | 8.3x | 1M | |
| gemini-2.5-flash-image | $0.300 | $2.500 | 8.3x | 32K | |
| gemini-3.1-flash-lite | $0.250 | $1.500 | 6.0x | 1M | |
| gemini-3.1-flash-lite-preview | $0.250 | $1.500 | 6.0x | 1M | |
| gpt-3.5-turbo | $0.500 | $1.500 | 3.0x | 16K | openai |
| gpt-3.5-turbo-0613 | $1.000 | $2.000 | 2.0x | 4K | openai |
| gpt-3.5-turbo-16k | $3.000 | $4.000 | 1.3x | 16K | openai |
| gpt-3.5-turbo-instruct | $1.500 | $2.000 | 1.3x | 4K | openai |
π΄ The Trap β Reasoning Models & High Output/Input Ratio
Section titled βπ΄ The Trap β Reasoning Models & High Output/Input RatioβThese models look cheap on input but the output costs 3.5β4.3x more. One long reasoning trace burns through tokens fast. This is the most likely cause of your surprise β¬2β3 DeepSeek bills.
| Model | Input/M | Output/M | Out/In | Ctx | Provider |
|---|---|---|---|---|---|
| deepseek-r1-0528 | $0.500 | $2.150 | 4.3x | 163K | deepseek |
| deepseek-r1 | $0.700 | $2.500 | 3.6x | 163K | deepseek |
| qwen3-30b-a3b-thinking-2507 | $0.080 | $0.400 | 5.0x | 131K | qwen |
| qwen3-next-80b-a3b-thinking | $0.098 | $0.780 | 8.0x | 262K | qwen |
| qwen3-235b-a22b-thinking-2507 | $0.150 | $1.500 | 10.0x | 262K | qwen |
| qwen3-vl-8b-thinking | $0.117 | $1.365 | 11.7x | 256K | qwen |
| qwen3-vl-30b-a3b-thinking | $0.130 | $1.560 | 12.0x | 131K | qwen |
| qwen3-vl-235b-a22b-thinking | $0.260 | $2.600 | 10.0x | 131K | qwen |
| qwen3-max-thinking | $0.780 | $3.900 | 5.0x | 262K | qwen |
| o3-mini | $1.100 | $4.400 | 4.0x | 200K | openai |
| o4-mini | $1.100 | $4.400 | 4.0x | 200K | openai |
| o3 | $2.000 | $8.000 | 4.0x | 200K | openai |
| o1 | $15.000 | $60.000 | 4.0x | 200K | openai |
π Anthropic / Claude Models
Section titled βπ Anthropic / Claude ModelsβConsistent 5x output-to-input ratio across the board. Clean pricing, no surprises β but premium absolute prices.
| Model | Input/M | Output/M | Out/In | Ctx |
|---|---|---|---|---|
| claude-3-haiku | $0.250 | $1.250 | 5.0x | 200K |
| claude-3.5-haiku | $0.800 | $4.000 | 5.0x | 200K |
| claude-haiku-4.5 | $1.000 | $5.000 | 5.0x | 200K |
| claude-sonnet-4.6 | $3.000 | $15.000 | 5.0x | 1M |
| claude-sonnet-4.5 | $3.000 | $15.000 | 5.0x | 1M |
| claude-sonnet-4 | $3.000 | $15.000 | 5.0x | 1M |
| claude-opus-4.8 | $5.000 | $25.000 | 5.0x | 1M |
| claude-opus-4.7 | $5.000 | $25.000 | 5.0x | 1M |
| claude-opus-4.6 | $5.000 | $25.000 | 5.0x | 1M |
| claude-opus-4 | $15.000 | $75.000 | 5.0x | 200K |
| claude-opus-4.1 | $15.000 | $75.000 | 5.0x | 200K |
| claude-opus-4.8-fast β‘ | $10.000 | $50.000 | 5.0x | 1M |
| claude-opus-4.7-fast β‘ | $30.000 | $150.000 | 5.0x | 1M |
| claude-opus-4.6-fast β‘ | $30.000 | $150.000 | 5.0x | 1M |
β‘ = fast mode variant, 2Γ the regular price for higher throughput
π OpenAI Flagship Models
Section titled βπ OpenAI Flagship Modelsβ| Model | Input/M | Output/M | Out/In | Ctx |
|---|---|---|---|---|
| gpt-4.1-nano | $0.100 | $0.400 | 4.0x | 1M |
| gpt-4o-mini | $0.150 | $0.600 | 4.0x | 128K |
| gpt-4.1-mini | $0.400 | $1.600 | 4.0x | 1M |
| gpt-4.1 | $2.000 | $8.000 | 4.0x | 1M |
| gpt-4o | $2.500 | $10.000 | 4.0x | 128K |
| gpt-4o-2024-08-06 | $2.500 | $10.000 | 4.0x | 128K |
| gpt-4o-search-preview | $2.500 | $10.000 | 4.0x | 128K |
| gpt-5-nano | $0.050 | $0.400 | 8.0x | 400K |
| gpt-5-mini | $0.250 | $2.000 | 8.0x | 400K |
| gpt-5 | $1.250 | $10.000 | 8.0x | 400K |
| gpt-5-chat | $1.250 | $10.000 | 8.0x | 128K |
| gpt-5.1 | $1.250 | $10.000 | 8.0x | 400K |
| gpt-5.4 | $2.500 | $15.000 | 6.0x | 1M |
| gpt-5.5 | $5.000 | $30.000 | 6.0x | 1M |
| gpt-5-pro | $15.000 | $120.000 | 8.0x | 400K |
| gpt-5.2 | $1.750 | $14.000 | 8.0x | 400K |
| gpt-5.3 | $1.750 | $14.000 | 8.0x | 400K |
| gpt-5.2-pro | $21.000 | $168.000 | 8.0x | 400K |
| gpt-5.4-pro | $30.000 | $180.000 | 6.0x | 1M |
| gpt-5.5-pro | $30.000 | $180.000 | 6.0x | 1M |
| gpt-4-turbo | $10.000 | $30.000 | 3.0x | 128K |
| gpt-4 | $30.000 | $60.000 | 2.0x | 8K |
| gpt-chat-latest | $5.000 | $30.000 | 6.0x | 400K |
π Google Gemini Models
Section titled βπ Google Gemini Modelsβ| Model | Input/M | Output/M | Out/In | Ctx |
|---|---|---|---|---|
| gemini-2.0-flash-lite-001 | $0.075 | $0.300 | 4.0x | 1M |
| gemini-2.0-flash-001 | $0.100 | $0.400 | 4.0x | 1M |
| gemini-2.5-flash-lite | $0.100 | $0.400 | 4.0x | 1M |
| gemini-2.5-flash | $0.300 | $2.500 | 8.3x | 1M |
| gemini-3.1-flash-lite | $0.250 | $1.500 | 6.0x | 1M |
| gemini-2.5-pro | $1.250 | $10.000 | 8.0x | 1M |
| gemini-3.1-pro-preview | $2.000 | $12.000 | 6.0x | 1M |
| gemini-3.5-flash | $1.500 | $9.000 | 6.0x | 1M |
| gemini-3-flash-preview | $0.500 | $3.000 | 6.0x | 1M |
| gemini-3.1-flash-image-preview | $0.500 | $3.000 | 6.0x | 131K |
π‘ Cost Scenario: 10-Turn Conversation
Section titled βπ‘ Cost Scenario: 10-Turn ConversationβAssumptions: ~5K context tokens per turn (full conversation history), ~1K output tokens per turn. Total: ~50K input + ~10K output.
This is the realistic Hermes session pattern β each turn sends the full context window.
| Model | Input Cost | Output Cost | Total |
|---|---|---|---|
| owl-alpha | FREE | FREE | β¬0.00 |
| deepseek-v4-flash | $0.005 | $0.002 | β¬0.007 |
| qwen3-235b-a22b-2507 | $0.004 | $0.001 | β¬0.005 |
| gpt-4.1-nano | $0.005 | $0.004 | β¬0.009 |
| deepseek-v3.2 | $0.013 | $0.004 | β¬0.016 |
| qwen3.6-35b-a3b | $0.007 | $0.010 | β¬0.017 |
| deepseek-chat-v3-0324 | $0.010 | $0.008 | β¬0.018 |
| gpt-4o-mini | $0.008 | $0.006 | β¬0.014 |
| gpt-5-mini | $0.013 | $0.020 | β¬0.033 |
| deepseek-r1-0528 | $0.025 | $0.022 | β¬0.047 |
| claude-3-haiku | $0.013 | $0.013 | β¬0.025 |
| deepseek-r1 | $0.035 | $0.025 | β¬0.060 |
| claude-haiku-4.5 | $0.050 | $0.050 | β¬0.10 |
| gpt-4.1 | $0.100 | $0.080 | β¬0.18 |
| gpt-4o | $0.125 | $0.100 | β¬0.23 |
| o3-mini | $0.055 | $0.044 | β¬0.10 |
| claude-sonnet-4.5 | $0.150 | $0.150 | β¬0.30 |
| gpt-5 | $0.063 | $0.100 | β¬0.16 |
| claude-opus-4.8 | $0.250 | $0.250 | β¬0.50 |
Key Takeaways
Section titled βKey Takeawaysβ- The out/in ratio is the hidden cost. Models like o3-mini and deepseek-r1 have 4x+ output multipliers. A 2K-token response at $4.40/M costs $0.009 β multiply that by 10+ turns with long reasoning traces and you hit β¬2β3 fast.
- Context window = compounding cost. Hermes sends full conversation history each turn. At 10K context Γ 20 turns = 200K input tokens. R1 input at $0.70/M = $0.14 for input alone before any output.
- Free models are legitimately free. owl-alpha at $0/M means your Hermes sessions cost nothing in LLM fees.
- Claude pricing is predictable (always 5x output) but premium. Sonnet 4.5 for a heavy session runs ~β¬0.30, Opus ~β¬0.50.
- The β¬2β3 mystery solved: Almost certainly deepseek-r1 or deepseek-chat-v3-0324 with verbose output, compounding over multiple turns with full context. Not a bug β just the hidden cost of reasoning models.