Density over Volume: Validation Density as the Binding Constraint on Specialist Post-Training
Specialist post-training scales corpus volume. We argue the real binding constraint is validation density, and report a compression result that tests it.
A benchmark tells you how a model performs on a curated task, in isolation, with everything around it held still. It cannot tell you what happens when the same model sits inside a working evidentiary pipeline — where retrieval decides what it reads, serving parameters decide whether it finishes, label construction decides what counts as correct, professional doctrine decides what a good answer even looks like, and a domain expert decides whether the output stands.
That is the stage after benchmarks, and it is the stage where deployed systems actually fail. RedTorch, Inc.® operates an environment where it can be measured.
These tracks cover selected questions arising from systems already operating inside our investigative platform. They are not a complete description of the platform, and not a complete inventory of current results.
Model behaviour in real evidentiary work is produced jointly by the model and by eight other things. The platform captures signals across all nine layers, with completeness measured separately rather than assumed — so some layers can be held fixed, others varied deliberately, and the places where attribution is still incomplete are identified rather than papered over. Where the instrumentation has holes, the holes are published; one of the five tracks exists to measure them.
In a benchmark you can vary one of these. Here you can vary any of them — and measure, rather than assume, which one was responsible.
We did not arrive at this position from theory. We arrived at it by being wrong, repeatedly, in the same direction — a model appeared to behave one way, and the cause turned out to sit in a layer around it.
An output budget silently truncated roughly a quarter of one extractor's valid output. Every recall number taken before it was raised understated the model.
Read the track →A compression run lost recall. Three architectural interventions moved it nowhere; changing how unproposed spans were weighted in the loss recovered it, and then some.
Read the track →Scoping a subgraph by "both endpoints came from this document" pulled in edges authored outside the intended scope. The graph arm looked richer while measuring a wider slice than the one under test — only per-producer provenance could tell them apart.
Read the track →A coverage catalogue counted note files instead of what experts had actually recorded, and reported a figure that was 74% false positives. Every count in it was accurate. An analyst caught it; no telemetry could have.
Read the track →None of those are discoverable from a benchmark score, and none of them are model problems. They are the reason this environment exists, and the reason we publish the failures alongside the results.
Each track states a position, reports what has been measured, and names the experiment that would settle it. Four are active with results in hand or measurement under way. One is proposed: the substrate exists, the experiment has not been run. All five are drafts and say so on the page.
Specialist post-training scales corpus volume. We argue the real binding constraint is validation density, and report a compression result that tests it.
Analysts spend most of their judgement ruling things out, and pipelines discard all of it. We treat expert rejections as typed, weighted hard negatives.
Agent autonomy is usually set in policy prose no runtime can check. We encode investigative doctrine as schema and report what enforcement actually buys.
Provenance is a profile, not a property. Any claim of full lineage must name its altitude and its weakest link — so we name ours, and decline the big one.
Vector retrieval finds similarity, not reasoning relevance. We condition retrieval on an adjudicated case graph — and on why scoping it needs provenance.
There is no tier list here on purpose. What a collaboration looks like depends on what you bring and what you want out of it, and we would rather design that with you once than publish a menu that fits nobody.
Tell us which track interests you and what you would want to run. If none of the five is quite it but the environment is, say that instead — the tracks are where our questions happen to be, not the boundary of what can be measured here.
Propose a collaborationThe environment is not a research rig. It is the platform our analysts work cases on every day, which is why the evidence in it is real and why the constraints are the ones deployment actually imposes. The platform itself is described under Technology; data and corpus arrangements under Licensing; the compliance posture that makes any of it possible under Compliance.