Thousand Brains v4 visual physics prediction
loading...

MPA record (PRIMARY)

no record yet

Visualise the world model

61 NVIDIA-FleX physics scenes, 3-frame → 1-frame prediction. Primary metric: Motion-Pixel Accuracy (MPA) — per-pixel accuracy restricted to the pixels that actually changed between frame_t+2 and the target. Strips the static-background "free score" baked into WMS. Models below the linear-extrapolation MPA floor are dimmed.

Live command center2 GPUs · 2 tracks

Pending evaluations

none

GPU hardware

idxnamemem usedutilwatts

Diversity gauges

Family share - target 40 / 30 / 30

Both tracks pick a family per cycle via the largest-deficit selector; champion ablates the top model in its family, exploration designs a fresh model in its family.

Conventional backprop share - cap 50%

A model is "conventional" iff it has feedforward_or_dense + cross_entropy_or_mse_loss + backprop_full_model AND no alternative-learning labels.
- of - completed runs 0% / 50%

Per-cycle family routing

Each cycle the 40/30/30 (BRAIN/FREE/MODERN) deficit selector picks one family per track. Champion ablates the top-MPA model in its family; exploration designs a fresh model in its family. Each track has its own last_*_family checkpoint for tie-breaks.

Beat-the-baseline race

Leaderboard0 models

# model family WMS vs linear (WMS) Motion % ↓ vs linear (MPA) per-scene SSIM PSNR params FPS Wh/MPA% archLabels
Reference baselines: -

Correlation labslice the data from every angle

Parallel coordinates - 9 dimensions, every model

Each line is a model. Drag through the cluster mentally: where do the high-MPA lines cross which axes? Hover a line to highlight.

Param count vs MPA - bubble = training Wh

More params doesn't buy MPA automatically. Smaller bubbles = leaner training.

Pareto frontier - efficiency vs MPA

The frontier line connects models that nothing else strictly dominates on (Wh, MPA). Living to the upper-left is winning.

Component-metric radar - top 5 by MPA

Each metric is normalised per-axis (0-1, higher is better). Look for shapes that bulge where the record holder is flat.

Per-scene heatmap (MPA primary, falls back to WMS)- models × 61 scenes

worse better -

Architecture label cloudtop 40 + bottom 40 by avg MPA

no labels yet
low avg MPA high avg MPA size = frequency

Label lifttop vs bottom MPA tertile

Top 12 - what helps

labelnliftavg MPA

Bottom 12 - what hurts

labelnliftavg MPA

Champion-track lineagesone per family-anchor base

The champion track now picks a family per cycle via the 40/30/30 (BRAIN/FREE/MODERN) deficit selector and ablates the top-MPA model in that family, so each family ends up with its own lineage. Versions in the same lineage share a base name; sibling lineages live in different families.
no lineages yet

Trajectory & efficiencyMPA primary

MPA over time vs baselines

Each dot is a finished evaluation. Step line traces the cumulative MPA record. Horizontal lines = MPA baselines (linear, static, quadratic, uniform-grey).

MPA-Lift histogram

Lift = model MPA% - linear-baseline MPA%. Anything ≤ 0 is worse than copy-frame extrapolation.

Wh per MPA% (training)

Lower bar = leaner model. Derived: training energy / (MPA × 100).

Inference FPS vs MPA

Generated 1024×1024 frames per second of wall-time during eval. Upper-right = accurate AND fast; lower-right = accurate but slow.

Inference FPS leaderboard

Top 24 evaluated models by raw throughput. Bar colour = family.

Forbidden patterns & negative-signal labels

    Raw briefings & learnings

    LEARNINGS.md (auto-generated digest, click to expand)
    loading...