OPEN RESEARCH

Model behaviour, measured inside the system that shapes it.

A benchmark tells you how a model performs on a curated task, in isolation, with everything around it held still. It cannot tell you what happens when the same model sits inside a working evidentiary pipeline — where retrieval decides what it reads, serving parameters decide whether it finishes, label construction decides what counts as correct, professional doctrine decides what a good answer even looks like, and a domain expert decides whether the output stands.

That is the stage after benchmarks, and it is the stage where deployed systems actually fail. RedTorch, Inc.® operates an environment where it can be measured.

These tracks cover selected questions arising from systems already operating inside our investigative platform. They are not a complete description of the platform, and not a complete inventory of current results.

WHAT YOU CAN MEASURE HERE

Nine layers, measured together.

Model behaviour in real evidentiary work is produced jointly by the model and by eight other things. The platform captures signals across all nine layers, with completeness measured separately rather than assumed — so some layers can be held fixed, others varied deliberately, and the places where attribution is still incomplete are identified rather than papered over. Where the instrumentation has holes, the holes are published; one of the five tracks exists to measure them.

  1. 01
    Source evidence
    Real privileged documents, as received — dense, inconsistent, partly scanned.
  2. 02
    Document and graph context
    What was retrieved and assembled, and by which rule.
  3. 03
    Model and checkpoint lineage
    Which weights, which quantization, which adapter, which template.
  4. 04
    Serving parameters
    Token budgets, caps, sampling, routing — the settings that quietly decide outcomes.
  5. 05
    Analyst decisions
    What a qualified expert accepted, dismissed, or ruled out.
  6. 06
    Accepted and rejected hypotheses
    Both halves of the judgement, kept as first-class signal.
  7. 07
    Professional doctrine
    The methodology standard analysts are held to, encoded as schema rather than prose.
  8. 08
    Operational failures
    Faults, timeouts, malformed output — recorded, not discarded.
  9. 09
    Autonomy boundaries
    What the system is permitted to do unattended, and on what evidence.

In a benchmark you can vary one of these. Here you can vary any of them — and measure, rather than assume, which one was responsible.

THE CASE FOR IT

Four times, the system explained the model.

We did not arrive at this position from theory. We arrived at it by being wrong, repeatedly, in the same direction — a model appeared to behave one way, and the cause turned out to sit in a layer around it.

None of those are discoverable from a benchmark score, and none of them are model problems. They are the reason this environment exists, and the reason we publish the failures alongside the results.

OPEN RESEARCH TRACKS

Five questions we are working on, and cannot finish alone.

Each track states a position, reports what has been measured, and names the experiment that would settle it. Four are active with results in hand or measurement under way. One is proposed: the substrate exists, the experiment has not been run. All five are drafts and say so on the page.

TRACK STAGE
ACTIVEmeasurement under way; results reported with their limits
PROPOSEDposition argued and substrate built; the deciding experiment has not been run
EVIDENCE
Measureda result exists, with a stated method and stated limitations
Observedfound in production and documented; not yet a controlled result
Implementedthe mechanism is built and running; completeness partly measured
Hypothesisargued and pre-specified; nothing has tested it
HOW THIS WORKS

We shape each program with the partner.

There is no tier list here on purpose. What a collaboration looks like depends on what you bring and what you want out of it, and we would rather design that with you once than publish a menu that fits nobody.

What RedTorch brings
  • Instrumentation spanning all nine layers, with coverage measured and gaps explicitly reported
  • Adjudicated ground truth from working investigations — acceptances, dismissals, ruled-out hypotheses, integrity findings
  • A doctrine substrate: a professional methodology taxonomy, embedded, with span-anchored groundings back to source
  • Existing harnesses — cross-validation folds, replayable prompts with lineage, paired arms behind one parameter
  • On-premises GPU capacity and a platform we control end to end, so an experiment can be run on privileged material without it leaving our custody
What partners have been asked for (examples, not requirements)
  • Training compute and a trainer, where we have the data density and no training service
  • Systems and kernel work, for the mechanisms we have specified and not built
  • A formal-methods eye, for enforcement properties we currently hold by convention
  • Sometimes only a second opinion on a measurement, which has already been worth more than compute

Tell us which track interests you and what you would want to run. If none of the five is quite it but the environment is, say that instead — the tracks are where our questions happen to be, not the boundary of what can be measured here.

Propose a collaboration

The environment is not a research rig. It is the platform our analysts work cases on every day, which is why the evidence in it is real and why the constraints are the ones deployment actually imposes. The platform itself is described under Technology; data and corpus arrangements under Licensing; the compliance posture that makes any of it possible under Compliance.