SIGMAAI

NIST AI RMF for evaluation teams

A practical playbook for operationalising the NIST AI Risk Management Framework inside the team that actually runs the evaluations — with concrete mapping steps for model assessment and the evidence you need to defend the result.

The NIST AI Risk Management Framework (AI RMF 1.0) is voluntary, but it is becoming the lingua franca of AI assurance. Regulators reference it, enterprise procurement teams ask about it, and internal audit functions increasingly score AI programmes against it. For evaluation teams, the RMF is less about new theory and more about a shared vocabulary for the work you already do: deciding what to test, how to test it, and what to keep on file when someone asks.

The four functions, in evaluation terms

The RMF is organised around four functions — Govern, Map, Measure, and Manage. Here is how each one lands on an evaluation team's desk.

Govern

Owns the rules of the game

Defines who signs off, what risk tolerances apply, and how evaluation findings flow into release decisions. For evaluation teams, this is your remit, escalation path, and the policy library you test against.

Map

Frames the system before testing

Captures the model's intended use, users, context, data lineage, and known limitations. The map is the brief your evaluation plan is written from — without it, test selection becomes guesswork.

Measure

Runs the evaluation itself

Quantitative and qualitative testing against the risks identified in Map: accuracy, robustness, bias, safety, and security. This is where evaluation teams spend most of their time.

Manage

Acts on what evaluation finds

Prioritises, mitigates, monitors, and retires AI risks over time. Evaluation feeds Manage by surfacing risk levels, trends across versions, and signals that a model needs to be re-reviewed or pulled.

Mapping steps for model evaluation

The Map function is where most evaluation programmes either compound or collapse. A weak map produces tests that pass but miss the real failure modes. Use the following sequence on every model that enters scope.

  1. Describe the intended use precisely. Write a one-paragraph statement covering who uses the model, in what workflow, with what inputs, and what decision the output drives. Anything outside that paragraph is out-of-scope use and gets its own risk entry.
  2. Identify the population the model touches. List the user segments, subjects of decisions, and any protected groups whose outcomes you are obliged to monitor under sector regulation (financial services, healthcare, employment, public sector).
  3. Document data lineage and provenance. Record the training data sources, licensing, collection windows, known gaps, and the pipeline that turns raw data into evaluation sets. Lineage is the single most-requested artefact in audits.
  4. Enumerate failure modes. For each intended use, write down the ways the model can plausibly fail — wrong answer, biased answer, unsafe answer, leaked data, manipulated output, degraded over time. This list drives test design in Measure.
  5. Assign a risk tier. Map each failure mode to severity and likelihood, then assign an overall risk tier the Govern function recognises. Tier determines how deep the evaluation goes and who has to sign off on the result.
  6. Lock the evaluation plan to the map. Every test in the plan should trace back to a failure mode, and every failure mode should have at least one test. Gaps in this matrix are the questions auditors will ask first.

Evidence collection that stands up later

The RMF does not prescribe a fixed evidence pack, but in practice you need enough material that a competent third party could reconstruct your conclusions six months later. Capture the following for each evaluation cycle.

  • Model card and system card. Versioned snapshot of the model, training data summary, intended use, and known limitations at the time of evaluation.
  • Evaluation plan. The risk-to-test matrix produced in Map, with the rationale for what is in scope and what is explicitly deferred.
  • Datasets and prompts. Hash-pinned evaluation sets, generation prompts, adversarial probes, and the provenance of each. Without this, results are not reproducible.
  • Run logs and metrics. Raw outputs, computed metrics, confidence intervals where applicable, and the environment the run was executed in.
  • Human review records. Reviewer instructions, calibration scores, inter-rater agreement, and the disposition of every flagged item. Human judgement only counts if it is auditable.
  • Findings and decisions. A signed summary linking results to risk tiers, the mitigation actions taken, and the named owner who accepted any residual risk.

Two practical rules make this evidence pack survive contact with reality. First, write to the pack as you go — retrofitting evidence after a release is the single most common audit finding. Second, version it alongside the model: each model version gets its own folder, its own hashes, and its own sign-off, even if 80% of the contents are unchanged.

Continuous evaluation under Manage

Evaluation does not end at release. The Manage function expects production monitoring that mirrors the categories tested pre-release — accuracy, bias, safety, and security drift. Set thresholds that, when crossed, trigger a re-run of the relevant Measure tests and a return trip through Govern.

Treat post-deployment incidents and user reports as inputs to the next Map cycle. A failure mode that surfaces in production should be added to the matrix, given a test, and re-evaluated before the next release. That feedback loop is what turns the RMF from a document into a working system.

Independent evaluation

Stand up an RMF-aligned evaluation programme

Sigma helps evaluation teams build the risk-to-test matrix, run the measurements, and assemble evidence packs that hold up under internal audit and external scrutiny.