﻿{"id":6190,"date":"2026-09-09T10:25:47","date_gmt":"2026-09-09T04:55:47","guid":{"rendered":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/?p=6190"},"modified":"2026-09-09T10:25:48","modified_gmt":"2026-09-09T04:55:48","slug":"continuous-outcome-evaluation-in-banking-how-ai-is-helping-bank-of-americas-competitors-continuously-evaluate-model-decision-outcomes","status":"publish","type":"post","link":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/artificial-intelligence\/continuous-outcome-evaluation-in-banking-how-ai-is-helping-bank-of-americas-competitors-continuously-evaluate-model-decision-outcomes.html","title":{"rendered":"Continuous Outcome Evaluation in Banking: How AI Is Helping Bank of America&#8217;s Competitors Continuously Evaluate Model &amp; Decision Outcomes"},"content":{"rendered":"<h4>Executive Summary<\/h4>\n<p>Continuous outcome evaluation is the discipline of checking whether a model&#8217;s real-world decisions \u2014 a fraud flag, a credit approval, an AI agent&#8217;s action \u2014 actually turn out to be correct, fair, and stable over time, rather than only checking the model once before it goes live. In banking, this discipline was formalized over a decade ago through the Federal Reserve and OCC&#8217;s SR 11-7 guidance, which required every model to be assessed for conceptual soundness, outcome analysis, and ongoing monitoring. What has changed in 2026 is the scale and speed AI demands of this process: generative and agentic AI systems drift, hallucinate, and change behavior far faster than the statistical models this guidance was originally written for.<\/p>\n<p>In a significant regulatory shift, the Federal Reserve, FDIC, and OCC rescinded SR 11-7, OCC 2011-12, and related guidance on April 17, 2026, replacing them with a more explicitly risk-based, principles-driven framework for model risk management \u2014 one that formally extends the same continuous-monitoring expectations to generative AI and agentic systems, even though those technologies are not yet named directly in the rule.<\/p>\n<p>Every Bank of America competitor examined here \u2014 JPMorgan Chase, Wells Fargo, Citigroup, Goldman Sachs, and Capital One \u2014 has built some version of continuous outcome evaluation into how it runs AI in production, ranging from Wells Fargo&#8217;s champion-challenger testing to Capital One&#8217;s AI agent that audits other AI agents. Bank of America&#8217;s own posture is grounded in a stated principle of human oversight and accountability for all outcomes, and in real-time monitoring of its credit and agentic AI systems, but \u2014 as with the bank&#8217;s broader AI-monitoring posture \u2014 it is less externally visible than some peers on the specific mechanics of how outcomes get continuously re-evaluated once a model is live.<\/p>\n<h4>What Continuous Outcome Evaluation Means in Banking<\/h4>\n<p>Model validation in banking has always rested on three pillars: conceptual soundness (is the model built on sound logic and data), outcome analysis (do its actual outputs turn out to be correct), and ongoing monitoring (does that accuracy hold up over time). Historically, this was a point-in-time discipline \u2014 a model was validated, approved, and then re-checked periodically, often annually.<\/p>\n<p>AI has broken that cadence. Traditional model risk frameworks assumed stable inputs, fixed logic, and outputs that could be documented once and reproduced consistently \u2014 assumptions that don&#8217;t hold for machine learning models that retrain on new data, or generative models whose behavior can shift between one prompt and the next. Model risk practice is consequently shifting from predominantly point-in-time controls toward continuous, risk-based practices spanning governance, development, validation, and monitoring together.<\/p>\n<p>The revised 2026 interagency guidance makes this explicit: performance drift, data drift, and model stability must now be tracked continuously, with detection thresholds mapped to how material a given model is to the business, rather than left to a calendar-based annual or semi-annual review cycle. The guidance also extends these expectations to third-party and vendor-built AI models, closing what had been a significant gap for banks that license rather than build their models.<\/p>\n<h4>Competitor Cases<\/h4>\n<h5>1. JPMorgan Chase<\/h5>\n<p><strong>How AI Is Helping in the Current Stage<br \/>\n<\/strong>JPMorgan Chase has extended continuous monitoring from portfolio risk into everyday AI operations. The bank uses AI to evaluate risk parameters constantly rather than only at quarterly review points, enabling continuous portfolio monitoring and real-time risk assessment instead of retrospective, point-in-time snapshots.<\/p>\n<p><strong>What Is the Impact<br \/>\n<\/strong>This continuous approach traces back to a hard lesson: JPMorgan&#8217;s 2012 &#8220;London Whale&#8221; trading loss of roughly $6 billion was rooted partly in a spreadsheet error that understated risk for months \u2014 a failure of outcome evaluation catching a bad result quickly. That history has shaped a bank-wide emphasis on evaluating the full range of possible outcomes continuously, not just relying on a model&#8217;s expected-case output.<\/p>\n<p><strong>Our POV<br \/>\n<\/strong>JPMorgan&#8217;s shift toward always-on, continuous risk evaluation \u2014 rather than periodic snapshots \u2014 is a direct, practical response to a real, expensive historical failure, which makes it a credible template rather than a theoretical best practice. The bank&#8217;s scale (over 450 AI use cases in production) means this discipline has to work across a very large and heterogeneous model inventory, which is the same challenge Bank of America will face as its own AI footprint grows.<\/p>\n<p><strong>Recommendations for Bank of America<\/strong><\/p>\n<ul>\n<li>Benchmark Bank of America&#8217;s current model review cadence against JPMorgan&#8217;s continuous, always-on risk evaluation approach, particularly for high-materiality models like credit and fraud scoring.<\/li>\n<li>Document a specific, named historical incident (internal or industry-wide) as the case for accelerating continuous outcome evaluation internally \u2014 JPMorgan&#8217;s London Whale narrative shows this framing helps sustain investment and urgency.<\/li>\n<\/ul>\n<h5>2. Wells Fargo<\/h5>\n<p><strong>How AI Is Helping in the Current Stage<br \/>\n<\/strong>Wells Fargo runs what its CIO has described as a champion-challenger framework: multiple competing model strategies are run continuously side-by-side in production, with their real-world performance compared over time to determine which one \u2014 the &#8220;champion&#8221; \u2014 is actually producing the best outcomes. On average, the bank deploys 50 to 60 models a year under this framework, with three independent groups \u2014 a frontline risk group, a model risk governance group, and an audit group \u2014 each building separate challenger models and independently checking results.<\/p>\n<p><strong>What Is the Impact<br \/>\n<\/strong>This structure gives Wells Fargo continuous, evidence-based outcome comparison rather than a single model&#8217;s self-reported accuracy. The bank pairs this with an explainability-first design principle \u2014 for example, its GPU-accelerated LIFE credit-decisioning algorithm was built specifically so that every output remains interpretable to regulators and customers even as the underlying models grow more complex, directly supporting continuous, ongoing outcome analysis of AI-driven credit models for drift and bias.<\/p>\n<p><strong>Our POV<br \/>\n<\/strong>The champion-challenger framework is the clearest, most operational example of continuous outcome evaluation in this entire peer set \u2014 it doesn&#8217;t just monitor one model for drift, it continuously tests whether a better model already exists, using genuinely independent teams to avoid the validator grading its own work. Combined with the explainability-first LIFE algorithm, Wells Fargo has built continuous evaluation and regulatory defensibility into the same pipeline rather than treating them as separate workstreams.<\/p>\n<p><strong>Recommendations for Bank of America<\/strong><\/p>\n<ul>\n<li>Evaluate adopting a formal champion-challenger structure for Bank of America&#8217;s highest-materiality models (credit decisioning, fraud, AML), with genuinely independent challenger-model teams rather than a single validation function.<\/li>\n<li>Pair any new AI-driven credit or risk model with an explainability requirement from day one, following Wells Fargo&#8217;s LIFE approach, rather than retrofitting interpretability after deployment.<\/li>\n<\/ul>\n<h5>3. Citigroup<\/h5>\n<p><strong>How AI Is Helping in the Current Stage<br \/>\n<\/strong>Citigroup has applied generative AI directly to one of the hardest outcome-evaluation problems in banking: keeping pace with new regulation. The bank used generative AI to read and assess the impact of 1,089 pages of new U.S. capital rules, compressing a task that would traditionally take a compliance team weeks into a far faster, continuously updatable review process.<\/p>\n<p><strong>What Is the Impact<br \/>\n<\/strong>Governance of this capability sits at board level: Citi&#8217;s Technology Committee holds explicit, documented oversight of generative AI and machine learning risks and benefits, meaning outcome evaluation for Citi&#8217;s AI systems is not just a model-validation-team function but a standing board responsibility \u2014 a materially higher bar of accountability than a purely technical monitoring pipeline.<\/p>\n<p><strong>Our POV<br \/>\n<\/strong>Citi&#8217;s approach shows that continuous outcome evaluation isn&#8217;t purely a technical exercise \u2014 regulatory-reading AI still needs its outputs checked against actual rule text and legal interpretation on an ongoing basis, since compliance rules keep changing. What stands out is governance: putting GenAI\/ML risk oversight explicitly in front of the board&#8217;s Technology Committee creates continuous accountability pressure that a purely technical monitoring dashboard cannot replicate on its own.<\/p>\n<p><strong>Recommendations for Bank of America<\/strong><\/p>\n<ul>\n<li>Consider elevating AI\/ML outcome-evaluation oversight to a named board or board-committee responsibility, following Citi&#8217;s model, rather than leaving it solely within a technology or model-risk function.<\/li>\n<li>Explore generative AI for continuously tracking and assessing the impact of new banking regulation on Bank of America&#8217;s own policies and models, as Citi has done, to keep outcome-evaluation criteria current as rules change.<\/li>\n<\/ul>\n<h5>4. Goldman Sachs<\/h5>\n<p><strong>How AI Is Helping in the Current Stage<br \/>\n<\/strong>Goldman Sachs&#8217; Model Risk Management (MRM) group \u2014 a multidisciplinary team of quantitative specialists spanning New York, Dallas, London, Birmingham, Warsaw, and Hong Kong \u2014 independently validates the firm&#8217;s AI and machine learning models for accuracy, explainability, design, and algorithmic robustness, separate from the teams that build them. AI models are evaluated and formally approved under this MRM framework, with bias detection and data lineage tracking built in before deployment.<\/p>\n<p><strong>What Is the Impact<br \/>\n<\/strong>Goldman has also centralized all internal AI activity onto a single, firewalled GS AI Platform specifically as a risk-management control: routing every use of AI, including its firmwide GS AI Assistant, through one governed gateway lets the bank enforce consistent policies on model validation, security, and auditability, rather than allowing business units to adopt inconsistent third-party tools with uneven oversight.<\/p>\n<p><strong>Our POV<br \/>\n<\/strong>Goldman&#8217;s centralized-platform strategy effectively makes continuous outcome evaluation structurally easier to enforce: if every AI interaction across the firm runs through one gateway, monitoring and re-evaluating outcomes becomes a platform-level capability rather than something each business unit has to build separately. This is a meaningfully different design choice than a federated approach, and it directly reduces the risk of ungoverned, unmonitored AI usage spreading inside the bank.<\/p>\n<p><strong>Recommendations for Bank of America<\/strong><\/p>\n<ul>\n<li>Assess whether Bank of America&#8217;s internal AI tools (Erica, ask MERRILL, ask PRIVATE BANK, and any future copilots) run through a single governed platform, following Goldman&#8217;s centralized model, to make continuous outcome monitoring a platform-wide capability rather than a per-tool responsibility.<\/li>\n<li>Adopt Goldman&#8217;s separation-of-duties principle explicitly \u2014 an MRM-style team independent from AI developers should own ongoing outcome validation for any high-materiality Bank of America AI system.<\/li>\n<\/ul>\n<h5>5. Capital One<\/h5>\n<p><strong>How AI Is Helping in the Current Stage<br \/>\n<\/strong>Capital One has built the most technically advanced continuous-evaluation architecture in this peer set. Its Machine Learning Experience (MLX) team is explicitly building an observability platform to monitor Capital One&#8217;s own generative AI and machine learning models in production, collecting metadata, metrics, and insights at scale specifically to track model and platform performance against compliance standards as models move from training into live, business-critical use.<\/p>\n<p><strong>What Is the Impact<br \/>\n<\/strong>In its production multi-agent system, Chat Concierge, Capital One embeds outcome evaluation directly into the AI architecture itself: after a customer-facing agent and a planning agent act, a third, independent evaluator agent audits what the first two did against Capital One&#8217;s own policies, and rejects and sends back any plan it judges non-compliant \u2014 creating an automated, continuous, iterative internal audit loop rather than relying only on downstream human log review.<\/p>\n<p><strong>Our POV<br \/>\n<\/strong>Capital One is the only bank examined here that has built continuous outcome evaluation directly into its AI&#8217;s runtime architecture, rather than treating it purely as an after-the-fact monitoring or periodic-review function. The evaluator-agent pattern is a genuinely novel approach: it turns &#8220;is this outcome acceptable&#8221; from a question answered later by a human reviewing logs into a question answered immediately, every time, by another AI system applying the bank&#8217;s own policies.<\/p>\n<p><strong>Recommendations for Bank of America<\/strong><\/p>\n<ul>\n<li>Pilot an evaluator-agent pattern \u2014 an independent AI system auditing the outputs of customer-facing or decisioning AI in real time \u2014 for Bank of America&#8217;s highest-volume automated workflows, starting with a contained, lower-risk use case.<\/li>\n<li>Invest in a dedicated internal team (comparable to Capital One&#8217;s MLX) whose sole mandate is building observability and continuous-evaluation tooling for Bank of America&#8217;s own AI\/ML models, rather than distributing this responsibility across individual product teams.<\/li>\n<\/ul>\n<h4>Comparative Snapshot: Continuous Outcome Evaluation<\/h4>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-6203\" src=\"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-content\/uploads\/2026\/09\/Banking-pic.png\" alt=\"\" width=\"1039\" height=\"645\" \/><\/p>\n<h4>Overall Point of View: Implications for Bank of America<\/h4>\n<p>Bank of America already states the right principle \u2014 human oversight, transparency, and accountability for all outcomes \u2014 and has real-time monitoring and control systems in place for its agentic AI and continuously-learning credit risk models. The gap this research surfaces is not principle but mechanism: peers have made continuous outcome evaluation an explicit, structural feature of how their AI runs, not just a policy statement layered on top of it.<\/p>\n<ul>\n<li>Make evaluation structural, not just procedural: Wells Fargo&#8217;s champion-challenger framework and Capital One&#8217;s evaluator-agent pattern both turn outcome evaluation into something the system does continuously and automatically, not something a review committee does periodically. Bank of America should identify at least one high-volume AI workflow to redesign this way.<\/li>\n<li>Centralize AI governance the way Goldman Sachs has: routing all internal AI tools \u2014 Erica, ask MERRILL, ask PRIVATE BANK, and future copilots \u2014 through a single governed platform would make continuous monitoring a shared capability rather than a per-product burden.<\/li>\n<li>Match the new regulatory bar directly: with SR 11-7 rescinded and replaced by an explicitly continuous, risk-tiered framework as of April 2026, Bank of America&#8217;s monitoring cadence for its highest-materiality models should be reviewed now against the new standard, rather than waiting for an examination finding to force the change.<\/li>\n<\/ul>\n<p>Recommended near-term focus areas: (1) pilot a champion-challenger or evaluator-agent pattern on one contained, high-volume AI workflow to build internal evidence before wider rollout; (2) assess whether Bank of America&#8217;s AI tools would benefit from a single governed platform akin to Goldman&#8217;s, to make continuous evaluation a shared, structural capability; and (3) formally map Bank of America&#8217;s current model monitoring cadence against the 2026 interagency guidance&#8217;s continuous, materiality-tiered monitoring expectations, closing any gaps ahead of the next examination cycle.<\/p>\n<h4>References<\/h4>\n<ul>\n<li>Klover.ai \u2014 Bank of America Uses AI Agents: 10 Ways to Use AI (2025) \u2014\u00a0<a href=\"https:\/\/www.klover.ai\/bank-of-america-uses-ai-agents-10-ways-to-use-ai-in-depth-analysis-2025\/\">https:\/\/www.klover.ai\/bank-of-america-uses-ai-agents-10-ways-to-use-ai-in-depth-analysis-2025\/<\/a><\/li>\n<li>VentureBeat \u2014 Wells Fargo CIO: AI and Machine Learning Will Move Financial Services Industry Forward \u2014 <a href=\"https:\/\/venturebeat.com\/ai\/wells-fargo-cio-ai-and-machine-learning-will-move-financial-services-ind\u2026\">https:\/\/venturebeat.com\/ai\/wells-fargo-cio-ai-and-machine-learning-will-move-financial-services-ind\u2026<\/a><\/li>\n<li>Klover.ai \u2014 Wells Fargo AI Strategy: Analysis of Dominance in Financial Services AI \u2014\u00a0<a href=\"https:\/\/www.klover.ai\/wells-fargo-ai-strategy-analysis-of-dominance-in-financial-services-ai\/\">https:\/\/www.klover.ai\/wells-fargo-ai-strategy-analysis-of-dominance-in-financial-services-ai\/<\/a><\/li>\n<li>McKinsey \u2014 Generative AI in Banking and Financial Services \u2014\u00a0<a href=\"https:\/\/www.mckinsey.com\/industries\/financial-services\/our-insights\/capturing-the-full-value-of-gen\u2026\">https:\/\/www.mckinsey.com\/industries\/financial-services\/our-insights\/capturing-the-full-value-of-gen\u2026<\/a><\/li>\n<li>Citigroup Inc. \u2014 Form DEF 14A, FY2024 (SEC proxy filing) \u2014 <a href=\"https:\/\/www.sec.gov\/Archives\/edgar\/data\/831001\/000114544324000041\/citi4284971-def14a.htm\">https:\/\/www.sec.gov\/Archives\/edgar\/data\/831001\/000114544324000041\/citi4284971-def14a.htm<\/a><\/li>\n<li>Indeed UK \u2014 Model Risk Model Validation Work, Goldman Sachs job listings \u2014 <a href=\"https:\/\/uk.indeed.com\/q-model-risk-model-validation-jobs.html\">https:\/\/uk.indeed.com\/q-model-risk-model-validation-jobs.html<\/a><\/li>\n<li>VentureBeat \u2014 How Capital One Built Production Multi-Agent AI Workflows \u2014\u00a0<a href=\"https:\/\/venturebeat.com\/ai\/how-capital-one-built-production-multi-agent-ai-workflows-to-power-enter\u2026\">https:\/\/venturebeat.com\/ai\/how-capital-one-built-production-multi-agent-ai-workflows-to-power-enter\u2026<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Executive Summary Continuous outcome evaluation is the discipline of checking whether a model&#8217;s real-world [&hellip;]<\/p>\n","protected":false},"author":914,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"inline_featured_image":false,"footnotes":""},"categories":[4,7],"tags":[78,1067,1069],"coauthors":[435],"class_list":["post-6190","post","type-post","status-publish","format-standard","hentry","category-artificial-intelligence","category-digital-experience","tag-ai","tag-banking","tag-countinousoutcomeevaluation"],"acf":[],"_links":{"self":[{"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/posts\/6190","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/users\/914"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/comments?post=6190"}],"version-history":[{"count":7,"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/posts\/6190\/revisions"}],"predecessor-version":[{"id":6193,"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/posts\/6190\/revisions\/6193"}],"wp:attachment":[{"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/media?parent=6190"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/categories?post=6190"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/tags?post=6190"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/blogs.infosys.com\/emerging-technology-solutions\/wp-json\/wp\/v2\/coauthors?post=6190"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}