Thousand Brains intelligence over memorization
loading…

Research dashboard

Every candidate is trained on tiny-math-textbooks, answers a fixed 100-question benchmark, and is scored by Claude. The scalar we chase is IQ. The scalar we ration is watts per IQ.

dataset: —
auto-refresh 60 s

GPU

idxnameUtilizationwattsmem usedtemp

Active training (per-track)

GPU 0 = champion (record-holder ablation). GPU 1 = exploration (45/45/10 family roll). Each subagent reads the OTHER slot's in-flight model so they don't accidentally run near-duplicate experiments in parallel.
GPU 0 · champion track
—
—
GPU 1 · exploration track
—
—
archLabels floor
—
at-floor models
archLabels median
—
median labels per model
modeltrackepochbatchlosslr elapsedlog age

Active training — loss curve

Live tail of every active model's train.log (one series per GPU slot), decimated to ~800 points each.

Primary goal leaders

The six metrics defined in CLAUDE.md. IQ is the scalar we chase; the rest are efficiency and capability constraints.
Record IQ (0–100, higher better)
—
—
Max correctness (0–1)
—
—
Max completeness (0–1)
—
—
Best training efficiency (W·h / IQ, lower better)
—
—
Best inference efficiency (kW·h / IQ / 1B, lower better)
—
—
Fastest inference (tokens / s, higher better)
—
—

Champion track — single-knob ablation lineage

GPU 0 is the champion track. Every run is a single-knob ablation of the current record holder. The next directory MUST be named —. Compounding ablations is how the record actually moves; exploration funds the next champion.
Record IQ
—
—
Days stale
—
—
Lineage size
—
—
Next free version
—
—
version modelId IQ params knob hint (notes excerpt)
Knob menu (pick exactly ONE on the next run)

Architecture-family diversity (exploration track)

GPU 1 rolls a family per run. Updated split: BRAIN 45 % · FREE 45 % · MODERN 10 %. MODERN was cut from 33 % because 27 prior runs averaged IQ ~5 — flagship-scale paradigms don't fit the 30 M / 12 h budget. Family colours are reused on every chart further down.
Next exploration roll
—
—
BRAIN runs (target 45 %)
—
—
FREE runs (target 45 %)
—
—
MODERN runs (target 10 %)
—
—
family count share target best model best IQ
Record IQ
—
—
Best lift vs baseline
—
—
Models evaluated
—
—
Active training
—
GPU 0 + GPU 1, 12 h cap
Most efficient (train)
—
W·h per IQ
Total project energy
—
across every finished run
Best arch
—
—

Visualisations

Per-model charts below are limited to the top 50 by IQ plus the baseline reference. —

Capability — how well each model reasons

Leaderboard IQ

Horizontal bars, sorted by IQ. Dashed line = baseline transformer, the bar every brain architecture must clear.

Primary goals — top 3 models

Radar of the six CLAUDE.md goals, normalised so outer = better on every axis.

Subscore breakdown per model

Grouped bars of the four judge subscores (50 knowledge Qs correctness/completeness + 50 mental-model Qs correctness/completeness). Sorted by IQ.

IQ delta vs baseline

Each brain architecture minus the baseline transformer's IQ. Positive bars = the architecture reasons better than the reference.

Reasoning quality — how the Claude judge grades answers

Knowledge vs mental-model

Per-model subscores. Points above the diagonal reason better than they recall.

Correctness vs completeness

The two axes the Claude judge scores each answer on. Upper-right = precise AND thorough.

Efficiency & scaling — the watt and parameter budgets

IQ vs parameters

Log-scaled params on x-axis. Points to the upper-left are compute-efficient.

IQ vs training energy

Watt-hours burned during training vs the resulting IQ. Upper-left is the efficient frontier.

Training efficiency

Watt-hours per point of IQ during training. Lower is better.

Inference efficiency

Watt-hours per IQ-point to generate 1 B tokens. Lower is better.

Inference throughput (tokens / s)

Tokens per second measured during eval. Higher = cheaper inference.

Pareto frontier — IQ vs inference cost

Each dot is an evaluated model. Upper-left dots define the efficient frontier — the models a new architecture needs to dominate.

Parameter-count buckets

Average and max IQ by model size. Helps spot over-/under-parameterised regimes.

Architecture & brain principles

Brain-principle lift

Top-tertile rate minus bottom-tertile rate for each principle. Positive = correlated with high IQ.

Arch Labels — top 40 by average IQ

Each bar is an architectural pattern (one entry from any model's archLabels). Bar height = average IQ across every model that wears the label. Models are encouraged to declare every pattern they implement, so a single model can contribute to dozens of bars.

Brain-inspired labels — IQ vs inference energy

Each dot is a brain-inspired architectural pattern (the 7 cortical principles plus every label that appears on any brain_inspired model). Y axis = average IQ across models wearing the label. X axis = average inference energy (W·h per 1B tokens, lower is better). Bubble size = number of models. Cyan = core cortical principle; blue = auxiliary brain-flavoured pattern. Upper-left points are the patterns to lean into.

Project history — where energy and time are going

Training wall-clock per model

Each point is a model: IQ on the X-axis, hours trained on the Y-axis. Red points were killed by the 12 h monitor cap.

Cumulative training energy

Running total of watt-hours burned across every finished training run. Tracks the project's real-world compute cost.

Leaderboard

Top 50 by IQ + baseline reference. —
# model arch labels IQ know / mental params W·h / IQ

Learnings digest

loading…