Model Cards
What if you could explore AI’s frontier through a set of trading cards? Compare models against each other, assess their safety evidence, and discover gaps in their reporting.
Join the waitlist ↗CAP is our subjective overall rating—not a benchmark score. About CAP
GPT-6 Astra1 of 28 · Tap a side card or swipe · Arrow keys to browse
GPT-6 Astra
ARC-AGI-3
99.9FrontierMath Tier 4 v2
97.6Terminal-Bench Science 0.1
64.6Agents’ Last Exam
59.3AutomationBench
41.4Advanced math and learning unfamiliar games are defining results of its launch. It also completes scientific workflows, though its business-process tests still leave substantial work unfinished.
Strengths & limitations
Strength
Interactive reasoning
Weakness
End-to-end automation
28 cards
The evidence landscape
What the published tests measure
541 distinct testsCounting & sources
Each named test has one primary category. Across all labs, a shared test counts once. A lab’s percentage uses its own inventory as the denominator. Safety includes misuse capabilities, alignment and oversight; these shares describe reporting, not model safety.
- GPT-6 Astra · Evaluation tables ↗
- GPT-6 Astra · PDF page 9–10 ↗
- GPT-6 Astra · Evaluation tables ↗
- GPT-6 Astra · PDF page 6–8 ↗
- GPT-6 Astra · Detailed benchmarks appendix; Thinking column ↗
- GPT-6 Astra · PDF page 5 ↗
- GPT-6 Astra · PDF page 3 ↗
- GPT-6 Astra · Evaluation tables ↗
- GPT-6 Astra · PDF page 6 ↗
- GPT-6 Astra · Evaluation tables ↗
- GPT-6 Astra · PDF page 4–5 ↗
- GPT-6 Astra · Detailed benchmarks, target GPT-5 / o3 column ↗
- GPT-6 Astra · PDF page 7 ↗
- GPT-6 Astra · What’s changed: AIME 2024 chart ↗
- GPT-6 Astra · PDF page 2–3 ↗
- GPT-6 Astra · PDF page 4–5 ↗
- GPT-5.6 Sol · Appendix ↗
- Claude Fable 5.1 · PDF page 188–190 ↗
- Claude Fable 5.1 · PDF page 173–174 ↗
- Claude Opus 4.8 · PDF page 220-221 ↗
- Claude Opus 4.8 · PDF page 205-207 ↗
- Claude Fable 5.1 · PDF page 295-296 ↗
- Claude Fable 5.1 · PDF page 135 ↗
- Claude Fable 5.1 · PDF p.178; supporting pp.167 ↗
- Claude Sonnet 5 · PDF p.199 ↗
- Kimi K3 · HTML table 2; target-model column ↗
- Kimi K3 · PDF page 26–27 ↗
- Grok 4.6 · PDF page 12 ↗
- Gemini 3.8 Flash · PDF page 3-4 ↗
- Gemini 3.8 Flash · PDF page 5 ↗
- Gemini 3.1 Pro · PDF page 4 ↗
- Muse Spark 1.3 · PDF page 1–4 ↗
- Muse Glimmer · HTML table 5; target-model column ↗
- Inkling · HTML table 1; target-model column ↗
- Inkling · HTML table 1; target-model column ↗
- Mistral Medium 3.5 · Benchmark graphic 1; Mistral Medium 3.5 column ↗
- Mistral Medium 3.5 · Benchmark graphic 2; Mistral Medium 3.5 column ↗
- DeepSeek V4-Pro · DeepSeek-V4-Pro-0813; not the Preview or Flash columns ↗
Data, coverage & definitions
Named tests and evaluation results use different counting units. A test counts once in the sharing analysis; the evaluation count can include its separately reported settings and subsets.
The inventory combines release tables with safety, capability and oversight assessments from system cards. The source pass now includes the long Anthropic and OpenAI reports and Kimi’s full technical report. All 23 canonical model entries have reconciled text-and-figure result counts. Named-test equivalence and testing conditions are assessed separately.
Dates are model release months, not document revision dates. Points represent publications or target-model counts, not a continuous model lineage. Axes stay fixed when filtering labs. A publication covering multiple models is represented by one word-count point, labeled with both names. Evaluation counts remain model-specific. Source texts include PDF and HTML publications.
Most shared tests counts distinct reporting labs, not matching test conditions. Explicit versions remain separate. Internal tests retain their lab identity. Testing-condition completeness, uncertainty estimates and independent replication still require a systematic review, so no percentages for those fields are shown.
| Model | Release | Document words | Reported results |
|---|---|---|---|
| o3 | Apr 2025 | 9,914 | 178 |
| GPT-5 | Aug 2025 | 18,968 | 170 |
| GPT-5.1 Thinking | Nov 2025 | 1,345 | 20 |
| GPT-5.2 Thinking | Dec 2025 | 6,433 | 72 |
| Gemini 3.1 Pro | Feb 2026 | 1,734 | 30 |
| GPT-5.3 Codex | Feb 2026 | 10,012 | 32 |
| GPT-5.4 | Mar 2026 | 10,692 | 87 |
| Claude Opus 4.7 | Apr 2026 | 60,962 | 482 |
| GPT-5.5 | Apr 2026 | 13,314 | 113 |
| Claude Opus 4.8 | May 2026 | 61,414 | 566 |
| Mistral Medium 3.5 | May 2026 | Unavailable | Unavailable |
| Claude Fable 5 | Jun 2026 | 81,280 | 111 |
| Claude Sonnet 5 | Jun 2026 | 34,007 | 243 |
| Claude Opus 5 | Jul 2026 | 48,174 | 357 |
| GPT-5.6 Sol | Jul 2026 | 20,386 | 155 |
| Inkling | Jul 2026 | 1,420 | 25 |
| Inkling-Small | Jul 2026 | 1,496 | 33 |
| Kimi K3 | Jul 2026 | 1,916 | 50 |
| DeepSeek V4-Pro | Aug 2026 | Unavailable | Unavailable |
| Grok 4.6 | Aug 2026 | 6,225 | 49 |
| Muse Glimmer | Aug 2026 | 2,107 | 30 |
| Muse Spark 1.2 | Aug 2026 | Unavailable | Unavailable |
| Claude Fable 5.1 | Sep 2026 | 55,405 | 142 |
| Claude Mythos 5.1 | Sep 2026 | 55,405 | 263 |
| Gemini 3.8 Flash | Sep 2026 | 1,216 | 20 |
| GPT-6 Astra | Sep 2026 | 28,004 | 265 |
| Muse Spark 1.3 | Sep 2026 | Unavailable | Unavailable |
About this set
Welcome to the first edition of the Free Systems Model Cards.
The explorer brings together roughly 476,000 words across 22 system and model cards, alongside supporting release pages and evaluation reports.
You can browse models, compare the benchmarks they share, see where labs are testing different things, trace claims back to their sources, and get a better sense of where the published evidence is genuinely comparable.
The physical deck captures the same snapshot of the frontier in collectible form.
All proceeds go back into the Free Systems community fund to support open, empirical research on AI and politics.
Sources, counting and comparison
Methodology
We use two evidence sets. Official capability sources determine which benchmark results can appear on a card. A narrower corpus of dedicated system and model cards supports comparisons of disclosure length and reporting practice.
Sources
Benchmark results come from first-party system cards, model cards, launch pages and evaluation reports. Each result is stored with its model, benchmark name and version, score, source location and reported test conditions.
Dedicated system and model cards form the canonical disclosure corpus. Launch pages and evaluation reports can support a capability result, but they do not count towards system-card length, evaluation totals or canonical-card overlap.
Benchmark results
A result appears only when the score can be located in an official source, assigned to the correct model and tied to the benchmark label and qualifiers reported beside it. Qualifiers such as tool access, effort level, subset and harness are retained because they can materially change a result.
These are published results, including some attributed to external evaluators. Free Systems has not rerun the evaluations.
Model comparisons
Comparisons follow a fixed order. We first look for the same benchmark with a sufficiently aligned setup. We then look for the same benchmark under different conditions, a related benchmark from the same family, and finally different benchmarks covering the same capability category.
The interface labels those relationships as direct comparison, qualified comparison, same benchmark with a different setup, related benchmark or different evidence. Only the first two support a score comparison. The others show how the labs chose to measure a capability; they do not establish which model is better.
Disclosure length
Disclosure length is the word count of the canonical system or model card. PDF counts use extracted document text; HTML counts use the saved article or model-card body and exclude navigation, scripts and page furniture. Each source is tied to a stored document hash so the count refers to a specific version.
A system card covering multiple models retains its full document count for each model. We do not divide a shared document into estimated model-level word counts.
Editorial fields and limitations
CAP is our subjective overall rating—not a benchmark score. We assign it editorially; it is not a standardized measurement or a calculated average of the bars. The creature artwork is illustrative.
The source material is uneven and changes over time. Labs use different benchmarks, versions, scaffolds and reporting conventions; many results cannot be reproduced from the published information alone. This is a working edition, and corrections will be recorded as the corpus is updated.
The physical deck
Series 1
Fall 2026
Join the waitlist Expected price: $15–$20 + shipping. No payment today.
Expected to ship within 2–3 weeks of orders opening. We’ll email you when the first print run is available.
All proceeds go back into the Free Systems community fund.
Questions? fsystems@stanford.edu




























