MPA record (PRIMARY)
no record yet
Visualise the world model
61 NVIDIA-FleX physics scenes, 3-frame → 1-frame prediction. Primary metric: Motion-Pixel Accuracy (MPA) — per-pixel accuracy restricted to the pixels that actually changed between frame_t+2 and the target. Strips the static-background "free score" baked into WMS. Models below the linear-extrapolation MPA floor are dimmed.
Live command center2 GPUs · 2 tracks
Pending evaluations
none
GPU hardware
| idx | name | mem used | util | watts |
|---|
Diversity gauges
Family share - target 40 / 30 / 30
Both tracks pick a family per cycle via the largest-deficit selector; champion ablates the top model in its family, exploration designs a fresh model in its family.
Conventional backprop share - cap 50%
A model is "conventional" iff it has feedforward_or_dense + cross_entropy_or_mse_loss + backprop_full_model AND no alternative-learning labels.
Per-cycle family routing
Each cycle the 40/30/30 (BRAIN/FREE/MODERN) deficit selector picks one family per track. Champion ablates the top-MPA model in its family; exploration designs a fresh model in its family. Each track has its own
last_*_family checkpoint for tie-breaks.Beat-the-baseline race
Leaderboard0 models
| # | model | family | WMS | vs linear (WMS) | Motion % ↓ | vs linear (MPA) | per-scene | SSIM | PSNR | params | FPS | Wh/MPA% | archLabels |
|---|
Reference baselines:
-
Correlation labslice the data from every angle
Parallel coordinates - 9 dimensions, every model
Each line is a model. Drag through the cluster mentally: where do the high-MPA lines cross which axes? Hover a line to highlight.
Param count vs MPA - bubble = training Wh
More params doesn't buy MPA automatically. Smaller bubbles = leaner training.
Pareto frontier - efficiency vs MPA
The frontier line connects models that nothing else strictly dominates on (Wh, MPA). Living to the upper-left is winning.
Component-metric radar - top 5 by MPA
Each metric is normalised per-axis (0-1, higher is better). Look for shapes that bulge where the record holder is flat.
Per-scene heatmap (MPA primary, falls back to WMS)- models × 61 scenes
worse
better
-
Architecture label cloudtop 40 + bottom 40 by avg MPA
no labels yet
low avg MPA
high avg MPA
size = frequency
Label lifttop vs bottom MPA tertile
Top 12 - what helps
| label | n | lift | avg MPA |
|---|
Bottom 12 - what hurts
| label | n | lift | avg MPA |
|---|
Champion-track lineagesone per family-anchor base
The champion track now picks a family per cycle via the 40/30/30 (BRAIN/FREE/MODERN) deficit selector and ablates the top-MPA model in that family, so each family ends up with its own lineage. Versions in the same lineage share a base name; sibling lineages live in different families.
no lineages yet
Trajectory & efficiencyMPA primary
MPA over time vs baselines
Each dot is a finished evaluation. Step line traces the cumulative MPA record. Horizontal lines = MPA baselines (linear, static, quadratic, uniform-grey).
MPA-Lift histogram
Lift = model MPA% - linear-baseline MPA%. Anything ≤ 0 is worse than copy-frame extrapolation.
Wh per MPA% (training)
Lower bar = leaner model. Derived: training energy / (MPA × 100).
Inference FPS vs MPA
Generated 1024×1024 frames per second of wall-time during eval. Upper-right = accurate AND fast; lower-right = accurate but slow.
Inference FPS leaderboard
Top 24 evaluated models by raw throughput. Bar colour = family.
Forbidden patterns & negative-signal labels
Raw briefings & learnings
LEARNINGS.md (auto-generated digest, click to expand)
loading...