AnantState
Decision horizon and model records

Decision horizon: how far ahead do you need to be right?

The first thing we ask is how far ahead you need to be right: tomorrow, next week, next quarter. That becomes the test the model must pass. We measure it at exactly that distance against the simplest possible guess, “nothing will change”, and write the result down, including where it loses.

In plain terms“Right about what, by when?” is the first question. The model is then graded on exactly that.

At a glance

  1. The horizon is declared

    A business input, not a tuning parameter.

  2. The null is “nothing changes”

    Not our previous release.

  3. Three verdicts, kept apart

    Passed, failed and not measured.

  4. A gate that refuses

    And records why.

A model that scores well and decides badly

Most model reporting has three defects.

  1. The horizon is chosen after the fact

    A model is scored at whatever distance flattered it. You deploy it and find it is useless at the distance you actually need.

  2. The baseline is another model

    Beating your own last version proves nothing. The real question is whether you beat assuming nothing changes.

  3. Failure is absent from the report

    A model record with no failure modes is a marketing document. You cannot take it to a validation committee.

We fix all three by making the horizon a declared input and the baseline a naive one.

Where the model wins, where it loses, and where it was never measured are three different facts. The horizon you declared says which one matters to you.

Four properties of a model that is allowed to serve

  • 1 · Declared per domain

    A domain states how far ahead it must be right, for example “20 steps ahead” (a step is one update of the model: a minute, an hour or a day, whichever suits the domain).

    That number is the objective the model is trained and judged against.

  • 2 · Reported per horizon

    Error at each horizon, never averaged away.

    A model that is excellent at one tick and useless at 20 is reported as exactly that. The shape of the curve is the information.

  • 3 · Compared to persistence

    Predict that the next state looks like the current one.

    The honest null. Where the ratio sits above 1, the model is worse than guessing “no change”, and the record says so.

  • 4 · Divergence is its own verdict

    Degenerate is not the same as inaccurate.

    “Retrain with more data” and “your data is malformed” are different instructions, so the gate refuses them in different words.

The promotion gate

What happens before a model is allowed to serve.

OrderCheckRefusal means
1Corpus provenanceThe data that trained this model is not attributable, or is too self-referential to certify
2DivergenceThe model’s output is degenerate, not merely inaccurate
3Beats the nullThe model does not beat persistence at the reported horizons
4Declared horizon measuredThe horizon you require was never actually measured

A refusal can be overridden by a human. When it is, the override is logged with the actor and the reason. The default posture is to refuse; using a bad model is a decision someone has to own. If the horizon you declared was never measured, the gate warns rather than silently passing: absence of evidence is recorded as absence of evidence, never converted into a pass.

The promotion gate refuses in a fixed order

  1. 1 · Corpus provenance
  2. 2 · Divergence
  3. 3 · Beats persistence
  4. 4 · Declared horizon measured
  5. Serve

The default posture is to refuse. A human can override, and the override is recorded with the actor and the reason.

What you get for every model

  • Per-horizon error, beside the persistence baseline at the same horizon
  • The furthest horizon at which the model still beats that baseline
  • The verdict: passed, failed or not measured, kept distinct
  • Corpus provenance: which teacher produced the training rows, how many are unattributed, and whether the corpus is multi-teacher
  • Promotion history: who promoted it, when, and whether the gate refused first
  • Drift and inventory context: how the serving model compares to its sibling revisions

This is what a validation committee asks for. It exists as a first-class surface, not a spreadsheet someone maintains by hand.

Model report card · illustrative shape, not a measured result

  • 1 stepPassed
  • 5 stepsPassed
  • 20 stepsFailed
  • 50 steps
    This horizon was never evaluated
    Not measured
Model errorThe simple guess: nothing changesFurthest distance ahead that still beats the simple guess: 5 steps
Each distance ahead is tested against the simple guess that nothing changes. Passed, failed and not measured stay three different results, and the furthest distance that still wins is recorded.

Why we publish the failures

“A vendor who shows you only their good horizons is asking you to trust their judgment about which horizons do not matter.”
  • You can deploy this without a leap of faith. The gate is a control you can point at.
  • You find out early when it is not right for you. A failure recorded on a small evaluation is cheaper than one discovered after an integration.

If the model does not beat persistence in your domain on your data, the honest outcome is that you are told, not that you are sold to.

The model record, on screen

The reference-world model record includes a horizon where the model loses. That is deliberate.

Prediction error by model on data the model never trained on, with a dashed baseline line: models that beat the baseline fall below it and those that do not rise above itFictional reference world: OgMart
Models are judged against a baseline on data they never trained on. The ones that lose are shown, not hidden.

Ask for the model record, including the losses

In a technical session we show the record for the reference world, then run an evaluation on a sample of your events.