Executive Summary
Continuous outcome evaluation is the discipline of checking whether a model’s real-world decisions — a fraud flag, a credit approval, an AI agent’s action — actually turn out to be correct, fair, and stable over time, rather than only checking the model once before it goes live. In banking, this discipline was formalized over a decade ago through the Federal Reserve and OCC’s SR 11-7 guidance, which required every model to be assessed for conceptual soundness, outcome analysis, and ongoing monitoring. What has changed in 2026 is the scale and speed AI demands of this process: generative and agentic AI systems drift, hallucinate, and change behavior far faster than the statistical models this guidance was originally written for.
In a significant regulatory shift, the Federal Reserve, FDIC, and OCC rescinded SR 11-7, OCC 2011-12, and related guidance on April 17, 2026, replacing them with a more explicitly risk-based, principles-driven framework for model risk management — one that formally extends the same continuous-monitoring expectations to generative AI and agentic systems, even though those technologies are not yet named directly in the rule.
Every Bank of America competitor examined here — JPMorgan Chase, Wells Fargo, Citigroup, Goldman Sachs, and Capital One — has built some version of continuous outcome evaluation into how it runs AI in production, ranging from Wells Fargo’s champion-challenger testing to Capital One’s AI agent that audits other AI agents. Bank of America’s own posture is grounded in a stated principle of human oversight and accountability for all outcomes, and in real-time monitoring of its credit and agentic AI systems, but — as with the bank’s broader AI-monitoring posture — it is less externally visible than some peers on the specific mechanics of how outcomes get continuously re-evaluated once a model is live.
What Continuous Outcome Evaluation Means in Banking
Model validation in banking has always rested on three pillars: conceptual soundness (is the model built on sound logic and data), outcome analysis (do its actual outputs turn out to be correct), and ongoing monitoring (does that accuracy hold up over time). Historically, this was a point-in-time discipline — a model was validated, approved, and then re-checked periodically, often annually.
AI has broken that cadence. Traditional model risk frameworks assumed stable inputs, fixed logic, and outputs that could be documented once and reproduced consistently — assumptions that don’t hold for machine learning models that retrain on new data, or generative models whose behavior can shift between one prompt and the next. Model risk practice is consequently shifting from predominantly point-in-time controls toward continuous, risk-based practices spanning governance, development, validation, and monitoring together.
The revised 2026 interagency guidance makes this explicit: performance drift, data drift, and model stability must now be tracked continuously, with detection thresholds mapped to how material a given model is to the business, rather than left to a calendar-based annual or semi-annual review cycle. The guidance also extends these expectations to third-party and vendor-built AI models, closing what had been a significant gap for banks that license rather than build their models.
Competitor Cases
1. JPMorgan Chase
How AI Is Helping in the Current Stage
JPMorgan Chase has extended continuous monitoring from portfolio risk into everyday AI operations. The bank uses AI to evaluate risk parameters constantly rather than only at quarterly review points, enabling continuous portfolio monitoring and real-time risk assessment instead of retrospective, point-in-time snapshots.
What Is the Impact
This continuous approach traces back to a hard lesson: JPMorgan’s 2012 “London Whale” trading loss of roughly $6 billion was rooted partly in a spreadsheet error that understated risk for months — a failure of outcome evaluation catching a bad result quickly. That history has shaped a bank-wide emphasis on evaluating the full range of possible outcomes continuously, not just relying on a model’s expected-case output.
Our POV
JPMorgan’s shift toward always-on, continuous risk evaluation — rather than periodic snapshots — is a direct, practical response to a real, expensive historical failure, which makes it a credible template rather than a theoretical best practice. The bank’s scale (over 450 AI use cases in production) means this discipline has to work across a very large and heterogeneous model inventory, which is the same challenge Bank of America will face as its own AI footprint grows.
Recommendations for Bank of America
- Benchmark Bank of America’s current model review cadence against JPMorgan’s continuous, always-on risk evaluation approach, particularly for high-materiality models like credit and fraud scoring.
- Document a specific, named historical incident (internal or industry-wide) as the case for accelerating continuous outcome evaluation internally — JPMorgan’s London Whale narrative shows this framing helps sustain investment and urgency.
2. Wells Fargo
How AI Is Helping in the Current Stage
Wells Fargo runs what its CIO has described as a champion-challenger framework: multiple competing model strategies are run continuously side-by-side in production, with their real-world performance compared over time to determine which one — the “champion” — is actually producing the best outcomes. On average, the bank deploys 50 to 60 models a year under this framework, with three independent groups — a frontline risk group, a model risk governance group, and an audit group — each building separate challenger models and independently checking results.
What Is the Impact
This structure gives Wells Fargo continuous, evidence-based outcome comparison rather than a single model’s self-reported accuracy. The bank pairs this with an explainability-first design principle — for example, its GPU-accelerated LIFE credit-decisioning algorithm was built specifically so that every output remains interpretable to regulators and customers even as the underlying models grow more complex, directly supporting continuous, ongoing outcome analysis of AI-driven credit models for drift and bias.
Our POV
The champion-challenger framework is the clearest, most operational example of continuous outcome evaluation in this entire peer set — it doesn’t just monitor one model for drift, it continuously tests whether a better model already exists, using genuinely independent teams to avoid the validator grading its own work. Combined with the explainability-first LIFE algorithm, Wells Fargo has built continuous evaluation and regulatory defensibility into the same pipeline rather than treating them as separate workstreams.
Recommendations for Bank of America
- Evaluate adopting a formal champion-challenger structure for Bank of America’s highest-materiality models (credit decisioning, fraud, AML), with genuinely independent challenger-model teams rather than a single validation function.
- Pair any new AI-driven credit or risk model with an explainability requirement from day one, following Wells Fargo’s LIFE approach, rather than retrofitting interpretability after deployment.
3. Citigroup
How AI Is Helping in the Current Stage
Citigroup has applied generative AI directly to one of the hardest outcome-evaluation problems in banking: keeping pace with new regulation. The bank used generative AI to read and assess the impact of 1,089 pages of new U.S. capital rules, compressing a task that would traditionally take a compliance team weeks into a far faster, continuously updatable review process.
What Is the Impact
Governance of this capability sits at board level: Citi’s Technology Committee holds explicit, documented oversight of generative AI and machine learning risks and benefits, meaning outcome evaluation for Citi’s AI systems is not just a model-validation-team function but a standing board responsibility — a materially higher bar of accountability than a purely technical monitoring pipeline.
Our POV
Citi’s approach shows that continuous outcome evaluation isn’t purely a technical exercise — regulatory-reading AI still needs its outputs checked against actual rule text and legal interpretation on an ongoing basis, since compliance rules keep changing. What stands out is governance: putting GenAI/ML risk oversight explicitly in front of the board’s Technology Committee creates continuous accountability pressure that a purely technical monitoring dashboard cannot replicate on its own.
Recommendations for Bank of America
- Consider elevating AI/ML outcome-evaluation oversight to a named board or board-committee responsibility, following Citi’s model, rather than leaving it solely within a technology or model-risk function.
- Explore generative AI for continuously tracking and assessing the impact of new banking regulation on Bank of America’s own policies and models, as Citi has done, to keep outcome-evaluation criteria current as rules change.
4. Goldman Sachs
How AI Is Helping in the Current Stage
Goldman Sachs’ Model Risk Management (MRM) group — a multidisciplinary team of quantitative specialists spanning New York, Dallas, London, Birmingham, Warsaw, and Hong Kong — independently validates the firm’s AI and machine learning models for accuracy, explainability, design, and algorithmic robustness, separate from the teams that build them. AI models are evaluated and formally approved under this MRM framework, with bias detection and data lineage tracking built in before deployment.
What Is the Impact
Goldman has also centralized all internal AI activity onto a single, firewalled GS AI Platform specifically as a risk-management control: routing every use of AI, including its firmwide GS AI Assistant, through one governed gateway lets the bank enforce consistent policies on model validation, security, and auditability, rather than allowing business units to adopt inconsistent third-party tools with uneven oversight.
Our POV
Goldman’s centralized-platform strategy effectively makes continuous outcome evaluation structurally easier to enforce: if every AI interaction across the firm runs through one gateway, monitoring and re-evaluating outcomes becomes a platform-level capability rather than something each business unit has to build separately. This is a meaningfully different design choice than a federated approach, and it directly reduces the risk of ungoverned, unmonitored AI usage spreading inside the bank.
Recommendations for Bank of America
- Assess whether Bank of America’s internal AI tools (Erica, ask MERRILL, ask PRIVATE BANK, and any future copilots) run through a single governed platform, following Goldman’s centralized model, to make continuous outcome monitoring a platform-wide capability rather than a per-tool responsibility.
- Adopt Goldman’s separation-of-duties principle explicitly — an MRM-style team independent from AI developers should own ongoing outcome validation for any high-materiality Bank of America AI system.
5. Capital One
How AI Is Helping in the Current Stage
Capital One has built the most technically advanced continuous-evaluation architecture in this peer set. Its Machine Learning Experience (MLX) team is explicitly building an observability platform to monitor Capital One’s own generative AI and machine learning models in production, collecting metadata, metrics, and insights at scale specifically to track model and platform performance against compliance standards as models move from training into live, business-critical use.
What Is the Impact
In its production multi-agent system, Chat Concierge, Capital One embeds outcome evaluation directly into the AI architecture itself: after a customer-facing agent and a planning agent act, a third, independent evaluator agent audits what the first two did against Capital One’s own policies, and rejects and sends back any plan it judges non-compliant — creating an automated, continuous, iterative internal audit loop rather than relying only on downstream human log review.
Our POV
Capital One is the only bank examined here that has built continuous outcome evaluation directly into its AI’s runtime architecture, rather than treating it purely as an after-the-fact monitoring or periodic-review function. The evaluator-agent pattern is a genuinely novel approach: it turns “is this outcome acceptable” from a question answered later by a human reviewing logs into a question answered immediately, every time, by another AI system applying the bank’s own policies.
Recommendations for Bank of America
- Pilot an evaluator-agent pattern — an independent AI system auditing the outputs of customer-facing or decisioning AI in real time — for Bank of America’s highest-volume automated workflows, starting with a contained, lower-risk use case.
- Invest in a dedicated internal team (comparable to Capital One’s MLX) whose sole mandate is building observability and continuous-evaluation tooling for Bank of America’s own AI/ML models, rather than distributing this responsibility across individual product teams.
Comparative Snapshot: Continuous Outcome Evaluation

Overall Point of View: Implications for Bank of America
Bank of America already states the right principle — human oversight, transparency, and accountability for all outcomes — and has real-time monitoring and control systems in place for its agentic AI and continuously-learning credit risk models. The gap this research surfaces is not principle but mechanism: peers have made continuous outcome evaluation an explicit, structural feature of how their AI runs, not just a policy statement layered on top of it.
- Make evaluation structural, not just procedural: Wells Fargo’s champion-challenger framework and Capital One’s evaluator-agent pattern both turn outcome evaluation into something the system does continuously and automatically, not something a review committee does periodically. Bank of America should identify at least one high-volume AI workflow to redesign this way.
- Centralize AI governance the way Goldman Sachs has: routing all internal AI tools — Erica, ask MERRILL, ask PRIVATE BANK, and future copilots — through a single governed platform would make continuous monitoring a shared capability rather than a per-product burden.
- Match the new regulatory bar directly: with SR 11-7 rescinded and replaced by an explicitly continuous, risk-tiered framework as of April 2026, Bank of America’s monitoring cadence for its highest-materiality models should be reviewed now against the new standard, rather than waiting for an examination finding to force the change.
Recommended near-term focus areas: (1) pilot a champion-challenger or evaluator-agent pattern on one contained, high-volume AI workflow to build internal evidence before wider rollout; (2) assess whether Bank of America’s AI tools would benefit from a single governed platform akin to Goldman’s, to make continuous evaluation a shared, structural capability; and (3) formally map Bank of America’s current model monitoring cadence against the 2026 interagency guidance’s continuous, materiality-tiered monitoring expectations, closing any gaps ahead of the next examination cycle.
References
- Klover.ai — Bank of America Uses AI Agents: 10 Ways to Use AI (2025) — https://www.klover.ai/bank-of-america-uses-ai-agents-10-ways-to-use-ai-in-depth-analysis-2025/
- VentureBeat — Wells Fargo CIO: AI and Machine Learning Will Move Financial Services Industry Forward — https://venturebeat.com/ai/wells-fargo-cio-ai-and-machine-learning-will-move-financial-services-ind…
- Klover.ai — Wells Fargo AI Strategy: Analysis of Dominance in Financial Services AI — https://www.klover.ai/wells-fargo-ai-strategy-analysis-of-dominance-in-financial-services-ai/
- McKinsey — Generative AI in Banking and Financial Services — https://www.mckinsey.com/industries/financial-services/our-insights/capturing-the-full-value-of-gen…
- Citigroup Inc. — Form DEF 14A, FY2024 (SEC proxy filing) — https://www.sec.gov/Archives/edgar/data/831001/000114544324000041/citi4284971-def14a.htm
- Indeed UK — Model Risk Model Validation Work, Goldman Sachs job listings — https://uk.indeed.com/q-model-risk-model-validation-jobs.html
- VentureBeat — How Capital One Built Production Multi-Agent AI Workflows — https://venturebeat.com/ai/how-capital-one-built-production-multi-agent-ai-workflows-to-power-enter…