Skip to content

Watching Runs in ScalarScope

A loss curve says how much a run learned. It does not say how: whether the student settled or kept wandering, whether the teachers’ judgements pulled the same way, or when a run regressed. aspire-si can record that as a training-dynamics export, and ScalarScope draws it and compares two runs.

Terminal window
aspire train --prompts data/prompts.json --geometry

or, in the config file:

training:
geometry_export: true
geometry_every: 1 # batches averaged into one export step
geometry_window: 4 # steps either side in the windowed measures

The trainer records every batch and writes geometry.json to the output directory when training ends. A run shorter than two recorded steps writes nothing and says so.

Terminal window
python examples/geometry_demo.py outputs/geometry-demo

This simulates two runs, one steady and one that regresses partway through, and writes both exports. It needs no model, GPU or API key. Open the two files on either side of ScalarScope’s Compare page.

Field What it holds
Trajectory The student’s last hidden layer, pooled over the unmasked tokens and the batch, projected onto its first two principal components and scaled so the farthest point is at distance 1.
Velocity The displacement per step.
Curvature The signed turning angle between steps, over pi. A turn counts in full only when the steps around it are as long as the run’s average step, so jitter after convergence does not read as instability.
Effective dimension The participation ratio of the hidden states in a window around the step: roughly how many directions the student is moving in.
Scalars Every evaluation dimension the teachers scored, from 0–10 scaled to 0–1.
Eigenvalues Per step, the spectrum of how the dimension scores vary together in a window, as fractions that sum to 1. A large first value means one direction explains the teachers’ judgements.
Professors One arrow per teacher: the direction in the 2-D state space along which that teacher’s score rises. A composite teacher gives one arrow per member.
Failures Steps where a dimension drops at least 0.1 below the median of the five steps before it. A dip lasting several steps is one failure.

The file is ScalarScope’s geometry format (schema 1.0). ScalarScope 3 reads every dimension; ScalarScope 2.0 reads only its five named dimensions.

from aspire.geometry import GeometryRecorder
recorder = GeometryRecorder(run_id="baseline", condition="socratic teacher", seed=42)
for batch in batches:
hidden = model(**batch, output_hidden_states=True).hidden_states[-1]
recorder.record(hidden, batch["attention_mask"], evaluations, teacher_name="socratic")
recorder.write("outputs/geometry.json", training_items=len(prompts), cycles=epochs)

record_step(state, scores, teacher_scores) takes a state vector and 0–10 scores directly, when your loop is not ASPIRE’s trainer.

The recorder keeps one pooled vector of the model’s hidden size per step. For long runs set geometry_every so several batches average into one step.

The trainer writes the export after every epoch as well as at the end, so a run stopped early (a deadline, Ctrl+C) keeps the epochs it finished.

The first real-model runs (a 1.5B student, 32B teachers, 3 epochs; see the run report) showed two things to keep in mind:

  • The per-step trajectory mostly shows which prompt each step drew. At batch 1, a step’s state is one prompt’s pooled hidden state. Two prompts sat about 10,000 times further apart than three epochs of training moved the student.
  • The scalars repeat each epoch. Epochs after the first replay the cached teacher scores, so the scalar curves and the dips show differences between prompts, not learning.

To see what training changed, replay a fixed set of exchanges through the base student and each epoch checkpoint and subtract each exchange’s base state. examples/pod-run/probe.py and examples/pod-run/drift.py do this and write a drift export: same format, prompt identity removed. In those runs the drift grew every epoch and pointed the same way for every exchange. It did not line up with the teacher’s scores, because 1.2.0 does not yet train the student toward better answers.

Exports are schema 1.1, and run_metadata states both of these facts:

Field Values
step_axis training_step: steps in training order. checkpoint_by_item: one block of the same items per checkpoint, in a fixed order, so the order within a block is not time.
checkpoints The number of blocks, present with checkpoint_by_item.
scalar_source live: every step scored afresh. replayed: epochs after the first reuse cached scores (the trainer with more than one epoch). fixed_per_item: each item’s scores repeat in every block.

ScalarScope uses them to withhold the comparisons that would read such steps as time, or replayed scores as new failures. GeometryRecorder(step_axis=..., checkpoints=..., scalar_source=...) sets them when you write your own export.