Research dashboard
Every candidate is trained on tiny-math-textbooks, answers a fixed 100-question benchmark, and is scored by Claude. The scalar we chase is IQ. The scalar we ration is watts per IQ.
dataset: —
auto-refresh 60 s
GPU
| idx | name | Utilization | watts | mem used | temp |
|---|
Active training (per-track)
GPU 0 = champion (record-holder ablation). GPU 1 =
exploration (45/45/10 family roll). Each subagent
reads the OTHER slot's in-flight model so they don't accidentally run
near-duplicate experiments in parallel.
GPU 0 · champion track
—
—
GPU 1 · exploration track
—
—
archLabels floor
—
at-floor models
archLabels median
—
median labels per model
| model | track | epoch | batch | loss | lr | elapsed | log age |
|---|
Active training — loss curve
Live tail of every active model's
train.log (one series per GPU slot), decimated to ~800 points each.Primary goal leaders
The six metrics defined in
CLAUDE.md. IQ is the scalar we chase; the rest are efficiency and capability constraints.
Record IQ (0–100, higher better)
—
—
Max correctness (0–1)
—
—
Max completeness (0–1)
—
—
Best training efficiency (W·h / IQ, lower better)
—
—
Best inference efficiency (kW·h / IQ / 1B, lower better)
—
—
Fastest inference (tokens / s, higher better)
—
—
Champion track — single-knob ablation lineage
GPU 0 is the champion track. Every run is a single-knob ablation of the
current record holder. The next directory MUST be named
—. Compounding ablations is how the
record actually moves; exploration funds the next champion.
Record IQ
—
—
Days stale
—
—
Lineage size
—
—
Next free version
—
—
| version | modelId | IQ | params | knob hint (notes excerpt) |
|---|
Knob menu (pick exactly ONE on the next run)
Architecture-family diversity (exploration track)
GPU 1 rolls a family per run. Updated split:
BRAIN 45 % · FREE 45 % ·
MODERN 10 %. MODERN was cut from 33 % because 27
prior runs averaged IQ ~5 — flagship-scale paradigms don't fit the
30 M / 12 h budget. Family colours are reused on every chart further down.
Next exploration roll
—
—
BRAIN runs (target 45 %)
—
—
FREE runs (target 45 %)
—
—
MODERN runs (target 10 %)
—
—
| family | count | share | target | best model | best IQ |
|---|
Record IQ
—
—
Best lift vs baseline
—
—
Models evaluated
—
—
Active training
—
GPU 0 + GPU 1, 12 h cap
Most efficient (train)
—
W·h per IQ
Total project energy
—
across every finished run
Best arch
—
—
Visualisations
Per-model charts below are limited to the top 50 by IQ plus the baseline reference. —
Capability — how well each model reasons
Leaderboard IQ
Horizontal bars, sorted by IQ. Dashed line = baseline transformer, the bar every brain architecture must clear.
Primary goals — top 3 models
Radar of the six CLAUDE.md goals, normalised so outer = better on every axis.
Subscore breakdown per model
Grouped bars of the four judge subscores (50 knowledge Qs correctness/completeness + 50 mental-model Qs correctness/completeness). Sorted by IQ.
IQ delta vs baseline
Each brain architecture minus the baseline transformer's IQ. Positive bars = the architecture reasons better than the reference.
Reasoning quality — how the Claude judge grades answers
Knowledge vs mental-model
Per-model subscores. Points above the diagonal reason better than they recall.
Correctness vs completeness
The two axes the Claude judge scores each answer on. Upper-right = precise AND thorough.
Efficiency & scaling — the watt and parameter budgets
IQ vs parameters
Log-scaled params on x-axis. Points to the upper-left are compute-efficient.
IQ vs training energy
Watt-hours burned during training vs the resulting IQ. Upper-left is the efficient frontier.
Training efficiency
Watt-hours per point of IQ during training. Lower is better.
Inference efficiency
Watt-hours per IQ-point to generate 1 B tokens. Lower is better.
Inference throughput (tokens / s)
Tokens per second measured during eval. Higher = cheaper inference.
Pareto frontier — IQ vs inference cost
Each dot is an evaluated model. Upper-left dots define the efficient frontier — the models a new architecture needs to dominate.
Parameter-count buckets
Average and max IQ by model size. Helps spot over-/under-parameterised regimes.
Architecture & brain principles
Brain-principle lift
Top-tertile rate minus bottom-tertile rate for each principle. Positive = correlated with high IQ.
Arch Labels — top 40 by average IQ
Each bar is an architectural pattern (one entry from any model's
archLabels). Bar height = average IQ across every model that wears the label. Models are encouraged to declare every pattern they implement, so a single model can contribute to dozens of bars.Brain-inspired labels — IQ vs inference energy
Each dot is a brain-inspired architectural pattern (the 7 cortical principles plus every label that appears on any
brain_inspired model). Y axis = average IQ across models wearing the label. X axis = average inference energy (W·h per 1B tokens, lower is better). Bubble size = number of models. Cyan = core cortical principle; blue = auxiliary brain-flavoured pattern. Upper-left points are the patterns to lean into.Project history — where energy and time are going
Training wall-clock per model
Each point is a model: IQ on the X-axis, hours trained on the Y-axis. Red points were killed by the 12 h monitor cap.
Cumulative training energy
Running total of watt-hours burned across every finished training run. Tracks the project's real-world compute cost.
Leaderboard
Top 50 by IQ + baseline reference. —
| # | model | arch labels | IQ | know / mental | params | W·h / IQ |
|---|
Learnings digest
loading…