Complete endpoints
Collection overview
Token entropy distribution
Reasoning tokens
Entropy across reasoning progress
Median, interquartile band, and 10th-90th percentile band
Fork alignment with double-newline boundaries
Literal \n\n boundaries in the reasoning region
Loading trajectory index
Preparing the endpoint browser.
Trajectory
-
-
Problem
Token entropy
Complete response
Response window
-
Measurement
Method
Teacher-forced entropy
Each response token is scored from the predictive distribution conditioned on the complete prompt and preceding response tokens. Entropy is reported in nats over the full model vocabulary.
Sampling replay
The secondary view reconstructs the collection policy: temperature 1.0, top-k 20, top-p 0.95, and output-token presence penalty 1.5. It does not draw new model samples.
Fork thresholds
Global thresholds are fit on train reasoning tokens and frozen for validation. Per-trace thresholds are independently measured within each reasoning trace. No local-maximum filter is applied.
Chunk alignment
Forks are compared with literal double-newline boundaries at exact, one-token, and two-token tolerances. Separator and solution tokens remain visible but are excluded from threshold fitting and alignment summaries.