OutFigure

Measurement infrastructure for learning products

What do your learners actually know, and how sure can you be?

An independent measurement layer for AI-powered and digital learning products. It keeps an auditable record of every observation, counts only evidence that is comparable, and reports learner state when the evidence supports it. When it does not, it says so.

The goal is to measure whether AI actually helped someone learn. That result does not exist yet; the evidence for it is being built one tested layer at a time.

See how the outcome audit works What has been tested

Evidence record

Synthetic example

Learner
synthetic_learner_a
Concept
DNS resolution
Sessions
1
  1. 09:02 Correct Counted
  2. 09:04 Incorrect Counted
  3. 09:07 Correct Not counted. Assistance was not reported
  4. 09:09 Hint shown Context for the next answer
  5. 09:10 Correct Not counted. A hint was shown before the answer
  6. 09:12 Correct Not counted. The item was already answered in this session
  7. 09:15 Correct Counted

2 of 3 counted responses correct

Insufficient evidence

forecast: null Not estimated. A summary needs 5 counted responses across 2 sessions.

From an event to a conclusion you can check

Every number traces back to the observations behind it, the rules that counted or excluded each one, and the version of those rules.

  1. 01

    Events

    Your backend sends answers, hints, feedback and interventions with pseudonymous IDs. Retries never count twice.

    Built
  2. 02

    Evidence

    Each response is counted or excluded for a stated reason: assisted, feedback visible, repeated, unscored.

    Built
  3. 03

    State

    Per learner and concept: what was observed, how recent, how thin. No mastery score and no guessed probability.

    Built
  4. 04

    Outcomes

    Predictions frozen before later, independent outcomes arrive, then compared honestly with simple baselines.

    In development

Begin with the data you already have

You should not need a full integration to find out whether this is useful. The outcome audit starts from an authorized, pseudonymized export of your existing assessment data.

  1. Readiness report

    Runs on your own machine. Validates the export, shows which context is missing, and states which analyses the data can and cannot support. No learner identifiers appear in the output.

    Available
  2. Baseline evaluation

    How well your current heuristic and simple baselines predict later responses, with calibration and uncertainty, on splits fixed before anyone looks at results.

    In development
  3. Reproducible report

    Every conclusion labeled descriptive, predictive or observational, reproducible from a manifest of the data and code versions.

    In development
  4. Continuous measurement

    If the audit is useful, the same measurement runs through the API as new events arrive.

    Local preview

Details of the audit Discuss an audit

What the system will not do

  1. Unknown stays unknown.

    Where evidence cannot support a value, the value is null with a reason. It is never a guessed percentage.

  2. Assisted success is not learning.

    Answers given with hints, outside help or unreported assistance are stored, but never counted as unaided evidence.

  3. Association is not cause.

    Each analysis names the kind of conclusion it can support. Observational data is never presented as proof that an intervention worked.

  4. Your data stays yours.

    No learner data is sent to language-model providers, and one customer's data is never used for another.

Where things stand

  • Working software, tested on synthetic data.
  • No customers or pilots yet.
  • No learning-outcome, prediction or efficacy results yet.
  • A hosted sandbox is in preparation.

Updated 15 September 2026