value rankings

Intelligence Per Dollar

Benchmark points per blended dollar across every credibly ranked model. Scores are min-max normalized across 253 models with 4+ independent benchmarks; cost assumes a typical 3:1 input:output token mix. Leader = 100.

This is not a quality ranking. #1 here means the most benchmark points per dollar, so a mid-scoring budget model will outrank a frontier model that costs 100x more. Check the Score and % of top score columns for raw capability, and use the task rankings when output quality is what compounds in your workflow.

Value leaderboard

#ModelScore% of top scoreIn / Out per 1MBlended $/1MValue
1inclusionAI: Ling 3.0 Flash BEST VALUEinclusionai/ling-3.0-flash39.339%$0.02 / $0.06$0.03100.0
2inclusionAI: Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl38.639%$0.02 / $0.06$0.0399.3
3Z.ai: GLM 5.3 Flashz-ai/glm-5.3-flash70.671%$0.04 / $0.14$0.0782.3
4Z.ai: GLM 5.3 Flash (batch)z-ai/glm-5.3-flash:batch70.671%$0.06 / $0.20$0.1059.6
5OpenAI: gpt-oss-120b (batch)openai/gpt-oss-120b:batch39.239%$0.03 / $0.14$0.0655.9
6OpenAI: gpt-oss-20bopenai/gpt-oss-20b24.324%$0.02 / $0.09$0.0454.1
7DeepSeek: DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash66.366%$0.04 / $0.29$0.1053.8
8OpenAI: GPT-6 Luna (batch)openai/gpt-6-luna:batch62.262%$0.05 / $0.25$0.1049.9
9OpenAI: gpt-oss-20b (batch)openai/gpt-oss-20b:batch24.324%$0.02 / $0.11$0.0542.3
10Mistral: Mistral Small 3mistralai/mistral-small-24b-instruct-250128.829%$0.05 / $0.08$0.0640.1
11Google: Gemma 4 26B A4B google/gemma-4-26b-a4b-it50.951%$0.07 / $0.23$0.1138.2
12Mistral: Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct59.660%$0.09 / $0.25$0.1336.0
13DeepSeek: DeepSeek V4.1 Flash (batch)deepseek/deepseek-v4.1-flash:batch66.366%$0.11 / $0.34$0.1731.6
14inclusionAI: Ling 3.0 Flash Fininclusionai/ling-3.0-flash-fin35.035%$0.06 / $0.18$0.0931.2
15Google: Gemma 4 31Bgoogle/gemma-4-31b-it53.754%$0.09 / $0.34$0.1528.2
16Google: Gemma 3 4Bgoogle/gemma-3-4b-it21.321%$0.05 / $0.10$0.0627.3
17OpenAI: GPT-6 Lunaopenai/gpt-6-luna62.262%$0.10 / $0.50$0.2024.9
18Upstage: Solar Pro 4upstage/solar-pro445.345%$0.09 / $0.36$0.1623.1
19OpenAI: GPT-5.6 Luna (batch)openai/gpt-5.6-luna:batch62.362%$0.10 / $0.60$0.2322.2
20Z.ai: GLM 4.7 Flashz-ai/glm-4.7-flash37.838%$0.06 / $0.40$0.1520.8
21Qwen: Qwen3 32Bqwen/qwen3-32b33.133%$0.08 / $0.28$0.1320.4
22OpenAI: GPT-5 Nano (batch)openai/gpt-5-nano:batch17.117%$0.03 / $0.20$0.0719.9
23Inception: Mercury 2.5inception/mercury-2.516.016%$0.04 / $0.15$0.0719.0
24OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch16.717%$0.05 / $0.20$0.0915.3
25DeepSeek: DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-exp57.758%$0.22 / $0.65$0.3214.3
26DeepSeek: DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro74.975%$0.35 / $0.70$0.4313.8
27StepFun: Step 3.5 Flashstepfun/step-3.5-flash24.524%$0.10 / $0.30$0.1513.1
28Qwen: Qwen3.5-9Bqwen/qwen3.5-9b18.318%$0.10 / $0.15$0.1113.0
29NVIDIA: Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning16.917%$0.08 / $0.20$0.1112.3
30DeepSeek: DeepSeek V3deepseek/deepseek-chat70.470%$0.32 / $0.89$0.4612.2

Budget champions : 80+ score, cheapest first

#ModelScore% of top scoreIn / Out per 1MBlended $/1MValue
1Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch94.294%$0.62 / $5.00$1.724.4
2OpenAI: o4 Mini Highopenai/o4-mini-high81.081%$1.10 / $4.40$1.933.4
3OpenAI: GPT-6 Sol (batch)openai/gpt-6-sol:batch81.281%$1.00 / $5.00$2.003.3
4Meta: Muse Spark 1.3meta/muse-spark-1.382.382%$1.25 / $4.25$2.003.3
5OpenAI: GPT-5.6 Sol (batch)openai/gpt-5.6-sol:batch80.280%$1.00 / $5.00$2.003.2
6OpenAI: GPT-5.4 (batch)openai/gpt-5.4:batch80.981%$1.25 / $7.50$2.812.3
7Google: Gemini 2.5 Progoogle/gemini-2.5-pro94.294%$1.25 / $10.00$3.442.2
8OpenAI: GPT-6 Solopenai/gpt-6-sol81.281%$2.00 / $10.00$4.001.6
9Anthropic: Claude Opus 5.5 (batch)anthropic/claude-opus-5.5:batch100.0100%$2.00 / $10.00$4.002.0
10OpenAI: GPT-5.6 Solopenai/gpt-5.6-sol80.280%$2.00 / $10.00$4.001.6

Free models with credible scores

Per-dollar math breaks at $0. These are simply the strongest free options:

Assumptions

Value = blended benchmark score divided by blended price per million tokens, indexed to the leader. A 3:1 input:output ratio fits most chat and RAG workloads; estimate your exact mix with the cost calculator. Scoring details in the methodology. Models with fewer than 4 independent benchmarks are excluded rather than guessed.

The Model Movers Report

One email every Friday, built from this site's own rankings: the current top five by benchmark score, every model released in the last seven days, and one note worked out from that week's numbers. You can unsubscribe from any issue with one click.