Model and data changes can be audited in production by tracing an evidence chain from approval and release through live behavior and corrective action. The record should identify what changed, connect pre-production checks to the deployed artifact, and preserve monitoring evidence that reviewers and responders can access. The cited material supports monitoring in production and retaining accessible monitoring data, but it does not prescribe a complete audit protocol.
Build a Traceable Change Record
A practical audit file can connect the following evidence. These are operational checks, not requirements stated in the cited sources.
| Stage | Evidence to retain | Audit question |
|---|---|---|
| Model change | Model or configuration identifier, change rationale, intended scope, review and approval records | Can the exact model change be identified? |
| Data change | Dataset or snapshot identifier, source, schema, transformations, labels or reference data, affected systems | Can the data change be followed through its downstream uses? |
| Validation | Test design, metrics, results, exceptions, and acceptance decision | Do the checks support the decision to release the change? |
| Deployment | Release time, deployed identifiers, relevant settings, and rollback status | Can the release be placed on the production timeline? |
| Production operation | System and component behavior, relevant model and data indicators, alerts, incidents, and interventions | What happened after deployment, and how was it handled? |
| Closeout | Acceptance decision, rollback record, follow-up actions, and unresolved risks | Is the final outcome and any remaining action traceable? |
Consistent identifiers are especially important. A model review can lose traceability when configuration changes but the model name remains the same. A data review can lose traceability when a dataset is updated without preserving the relevant snapshot or transformation history.
Check Before and After Deployment
A production audit should compare the released model or data state with the state that existed before the change. The comparison can use task-appropriate measures such as output quality, system errors, latency, component behavior, or relevant data characteristics. The appropriate measures depend on the workload; the cited material does not establish universal thresholds.
A reviewer can proceed by:
- Selecting a specific model or data release and reconstructing its approval, validation, deployment, and intervention timeline.
- Confirming that the test results, deployed artifact, and production records refer to the same version or data snapshot.
- Comparing pre-change and post-change behavior using criteria documented before deployment where possible.
- Examining both end-to-end results and the behavior of individual system components.
- Reviewing alerts, incidents, operator interventions, and follow-up decisions associated with the period after release.
- Recording contradictory evidence, missing records, and unresolved differences rather than treating missing documentation as proof that no change occurred.
The NIST AI RMF states that the functionality and behavior of the AI system and its components, as identified in its map function, are monitored when in production. This supports component-level monitoring rather than relying only on the final output.
Preserve Evidence for Production Review
The cited monitoring guidance says monitoring data should be retained and accessible to support timely detection, response, and post-incident analysis. In practice, accessibility means that an authorized reviewer can connect monitoring records to the relevant production event without relying on an unrecorded recollection.
A post-incident audit can use that evidence to determine which model or data version was active, when observable behavior changed, whether alerts were raised, what response followed, and whether the outcome was accepted, reversed, or escalated. The audit can establish what the available records show at the time of review; it cannot guarantee future performance or identify causes that were never recorded.
What Operators Must Still Confirm
Before treating this process as sufficient, operators must establish several matters that the cited passages do not specify:
- Which model and data changes require review, and who may approve them.
- Which metrics and thresholds trigger investigation or rollback.
- How long model versions, data snapshots, monitoring records, and audit logs are retained.
- Who may access production evidence and under what controls.
- How data lineage, privacy, security, contractual, and other applicable requirements affect the record.
- When a change is considered accepted, when it must be reversed, and who closes the audit.
- Which evidence must be produced for an internal review, external audit, incident investigation, or dispute.
A production audit is therefore an evidence trail rather than a declaration of correctness. Its reliability depends on consistent change records, observable system and component behavior, and retained monitoring data that remains available for later review.