skip to content
Conifer · Models

Models

Size · speed · intelligence.

The field on one axis. Cyan marks what Conifer runs locally.

Intelligence

Artificial Analysis Intelligence Index · higher is better · Artificial Analysis · 60 models

  1. 59.9
    Claude Fable 5: 59.9 index
  2. 55.7
    Claude Opus 4.8: 55.7 index
  3. 54.8
    GPT-5.5: 54.8 index
  4. 53.5
    Claude Opus 4.7: 53.5 index
  5. 44.3
    DeepSeek V4 Pro: 44.3 index
  6. 44.2
    Kimi K2.6: 44.2 index
  7. 40.8
    Claude Opus 4.5: 40.8 index
  8. 40
    GPT-5.4 mini: 40 index
  9. 39.6
    Gemini 3 Pro: 39.6 index
  10. 38.2
    GPT-5.4 nano: 38.2 index
  11. 36.9
    GPT-5.1 high: 36.9 index
  12. 36.4
    Claude Sonnet 4.5: 36.4 index
  13. 34.7
    GPT-5: 34.7 index
  14. 33.3
    Grok 4: 33.3 index
  15. 32.7
    Kimi K2 Thinking: 32.7 index
  16. 32
    DeepSeek V3.2: 32 index
  17. 31.7
    Qwen3 Max: 31.7 index
  18. 30.6
    Grok 4.1 Fast: 30.6 index
  19. 28.7
    GLM-4.6: 28.7 index
  20. 27.4
    Grok 4 Fast: 27.4 index
  21. 25.8
    Gemini 2.5 Pro: 25.8 index
  22. 23.8
    gpt-oss 120B high: 23.8 index
  23. 21.8
    Nova 2.0 Pro: 21.8 index
  24. 20.9
    Nova 2.0 Omni: 20.9 index
  25. 20.1
    DeepSeek R1 0528: 20.1 index
  26. 20.1
    Gemini 2.5 Flash: 20.1 index
  27. 20.1
    Qwen3.5 4B: 20.1 index, runs locally on Conifer
  28. 19.6
    Qwen3 235B 2507: 19.6 index
  29. 19.2
    Devstral 2: 19.2 index
  30. 18.2
    Nova 2.0 Lite: 18.2 index
  31. 17.9
    Qwen3 VL 32B: 17.9 index
  32. 17.7
    MiniMax M1: 17.7 index
  33. 17
    o1-preview: 17 index
  34. 16.5
    GLM-4.5 Air: 16.5 index
  35. 15.9
    Mistral Large 3: 15.9 index
  36. 15.4
    DeepSeek V3 0324: 15.4 index
  37. 14.9
    gpt-oss 20B: 14.9 index
  38. 14.8
    GPT-4.1 mini: 14.8 index
  39. 14.4
    Qwen3 30B A3B 2507: 14.4 index, runs locally on Conifer
  40. 14.3
    Llama 4 Maverick: 14.3 index
  41. 14.2
    Nemotron 3 Nano: 14.2 index
  42. 12.5
    Mistral Medium 3: 12.5 index
  43. 11.5
    Qwen3 32B: 11.5 index, runs locally on Conifer
  44. 11
    R1 Distill 32B: 11 index, runs locally on Conifer
  45. 10.4
    Qwen3 14B: 10.4 index, runs locally on Conifer
  46. 9.8
    R1 Distill 14B: 9.8 index, runs locally on Conifer
  47. 9.4
    Llama 3.3 70B: 9.4 index
  48. 8.4
    Qwen3 4B: 8.4 index, runs locally on Conifer
  49. 8.3
    Qwen3 8B: 8.3 index, runs locally on Conifer
  50. 7.6
    Llama 3.1 8B: 7.6 index, runs locally on Conifer
  51. 7.5
    Qwen2.5 32B: 7.5 index, runs locally on Conifer
  52. 7.4
    Gemma 3 27B: 7.4 index
  53. 7.1
    Qwen2.5 Coder 32B: 7.1 index, runs locally on Conifer
  54. 6.4
    R1 Distill Llama 8B: 6.4 index, runs locally on Conifer
  55. 5.5
    Gemma 3 12B: 5.5 index, runs locally on Conifer
  56. 4.9
    Phi-4: 4.9 index, runs locally on Conifer
  57. 4.5
    Qwen2.5 Coder 7B: 4.5 index, runs locally on Conifer
  58. 4.2
    Llama 3.2 3B: 4.2 index, runs locally on Conifer
  59. 2.6
    Qwen3 1.7B: 2.6 index, runs locally on Conifer
  60. 1.1
    Gemma 3 4B: 1.1 index, runs locally on Conifer
frontier & open runs locally on Coniferpublished figures · snapshot Jul 2026
how to read this

The whole Artificial Analysis field on one axis, and lit in cyan: the open models Conifer runs locally.

On the Intelligence composite the cloud giants tower; flip to GPQA, AIME or coding to watch the laptop-sized models punch far above their weight.

methodology & sources

Published figures, cross-checked against Artificial Analysis and each model's own technical report (snapshot Jul 2026— refreshed automatically from the live AA leaderboard).

Intelligence is the live Artificial Analysis Intelligence Index (a multi-eval composite on AA's current scale); capability scores are reasoning / thinking mode, no external tools, and benchmark versions vary by lab.

“Runs locally” marks the open weights in Conifer's catalogue that fit a personal machine; larger open models (Kimi, DeepSeek V3/V4, gpt-oss 120B, Qwen3 235B) sit in the field.

How far behind is one consumer GPU? Six to twelve months.

  • AA Intelligence Index6.3 mo
  • MMLU-Pro7.3 mo
  • GPQA-Diamond7.4 mo
  • LM Arena Elo12.4 mo
Average lag between the frontier and the best open model that fits one consumer GPU · Epoch AI, Aug 2025 (CC-BY) · source & method
source & method

Epoch AI, “Frontier AI performance becomes accessible on consumer hardware within a year” (Somala & Emberson, Aug 2025, CC-BY): the best open model that fits a single consumer GPU matches frontier scores from 6–12 months earlier — and it is gaining (+125 vs +80 Elo per year on LM Arena).

Their bar for “fits”: full weights, 4-bit quantized, in one card's VRAM — models to ~28B on an RTX 4090, ~40B on an RTX 5090. Epoch's caveat: small open models are likelier to be benchmark-tuned, so real-world lag may run somewhat longer.

Supported31 models
modelsizetok/sintelligence
Llama
Llama 3.2 1B1.2B29949.3
Llama 3.2 3B3.2B11863.4
Llama 3.1 8B8.0B4969.4
Hermes 3 · Llama 3.1 8B8.0B~4964.8
Qwen 3
Qwen 3 0.6B0.6B~29052.8
Qwen 3 1.7B1.7B~16562.6
Qwen 3 4B4.0B9073.0
Qwen 3 4B Instruct 25074.0B~90·
Qwen 3 8B8.2B4676.9
Qwen 3 14B14.8B~2581.1
Qwen 3 32B32.8B~1283.6
Qwen 3 30B A3B 250730.5B · 3B active·87.1
Qwen 3 Coder 30B A3B30.5B · 3B active··
Qwen 2.5
Qwen 2.5 0.5B0.5B28847.5
Qwen 2.5 1.5B1.5B16660.9
Qwen 2.5 3B3.1B8965.6
Qwen 2.5 7B7.6B5374.2
Qwen 2.5 14B14.7B~2779.7
Qwen 2.5 Coder 1.5B1.5B~16653.6
Qwen 2.5 Coder 7B7.6B~5368.0
Gemma
Gemma 2 2B2.6B13851.3
Gemma 2 9B9.2B~4171.3
Gemma 3 4B4.3B6659.6
Gemma 4 12B11.9B1574.5
DeepSeek
R1 Distill Qwen 1.5B1.8B~166·
R1 Distill Qwen 7B7.6B~53·
R1 Distill Llama 8B8.0B~49·
R1 Distill Qwen 14B14.8B~27·
R1 Distill Qwen 32B32.8B~12·
DeepSeek V2 Lite15.7B·55.7
Phi
Phi 3.5 Mini3.8B~7869.0
measurement notes

tok/s: decode, Apple M3 Max, Q4_K_M, 512-token prompt; ~ projected from a measured sibling. intelligence: MMLU 5-shot, as published. · marks not measured / not published.