﻿{"id":8320,"date":"2026-09-24T13:39:43","date_gmt":"2026-09-24T08:09:43","guid":{"rendered":"https:\/\/blogs.infosys.com\/digital-experience\/?p=8320"},"modified":"2026-09-24T13:39:43","modified_gmt":"2026-09-24T08:09:43","slug":"generative-ui-how-ai-agents-are-starting-to-build-their-own-interfaces","status":"publish","type":"post","link":"https:\/\/blogs.infosys.com\/digital-experience\/emerging-technologies\/generative-ui-how-ai-agents-are-starting-to-build-their-own-interfaces.html","title":{"rendered":"Generative UI: How AI Agents Are Starting to Build Their Own Interfaces"},"content":{"rendered":"<p>Most conversations with AI still happen through a single medium: text. You ask, it answers, in a scrolling wall of prose. But a growing share of real tasks \u2014 booking a flight, filling out a form, comparing products, exploring a dataset \u2014 are fundamentally\u00a0visual and interactive. Text is a bad interface for those.<\/p>\n<p>Generative UI (GenUI) is the emerging answer, instead of only replying with words, an AI agent generates, selects, or controls parts of the user interface at runtime \u2014 rendering the widget, form, or layout the moment actually calls for, rather than forcing everything through a chat bubble.<\/p>\n<p><strong>Why It Matters?<\/strong><br \/>\nText is a poor interface for structured tasks. A date picker, a comparison table, or a multi-step form is faster and clearer than the equivalent paragraph of prose. It closes the gap between &#8220;chatbot&#8221; and &#8220;app.&#8221;<\/p>\n<p>Instead of hand-coding a screen for every possible user intent, you build a component library and let the model assemble or select from it based on context. It reduces development overhead.\u00a0The interface adapts to what the user actually needs in the moment, rather than being fully predetermined by a developer months in advance.<\/p>\n<p>Search interest in the term has roughly doubled over the past year, and the framework names around it \u2014 Vercel AI SDK, CopilotKit, AG-UI \u2014 are seeing steadily rising, low-competition search volume, a sign that tooling is still catching up to demand.<\/p>\n<p><strong>How It Works?<\/strong><\/p>\n<p>Generative UI implementations generally fall into one of three\u00a0 architectural patterns, trading off control for flexibility.<\/p>\n<p><strong><em>1. Controlled Generative UI<\/em><\/strong><br \/>\nHigh control, low freedom. You pre-build a fixed set of components. The agent&#8217;s only job is to decide when to show one and what data to pass into it \u2014 typically via standard tool\/function calling. This is the safest, most predictable pattern, and the one most production apps use today for transactional flows (checkout, settings, bookings).<\/p>\n<p><strong><em>2. Declarative Generative UI<\/em><\/strong><br \/>\nShared control.\u00a0The agent returns a structured schema \u2014 JSON describing cards, lists, forms, layout \u2014 and the frontend renders it against a known component catalog. The agent can compose novel arrangements, but only from primitives you&#8217;ve already built, so output stays bounded.<\/p>\n<p>This is where specs like\u00a0A2UI\u00a0and\u00a0Open-JSON-UI\u00a0live.<\/p>\n<p><strong><em>3. Open-Ended Generative UI<\/em><\/strong><br \/>\nLow control, high freedom.\u00a0The agent returns an entire UI surface \u2014 raw HTML\/SVG\/Canvas, or a full iframe served by an external MCP server. The frontend becomes a sandboxed container that just displays whatever comes back.<\/p>\n<p>This is the most powerful pattern \u2014 capable of rendering diagrams, 3D scenes, D3 force layouts, or live simulations \u2014 but it requires careful sandboxing, since you&#8217;re rendering agent-generated code directly.<\/p>\n<p><strong>The Protocol Stack Underneath<\/strong><\/p>\n<p>Three open, complementary protocols have emerged, each addressing a different layer of the agentic stack:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignleft wp-image-8829\" src=\"https:\/\/blogs.infosys.com\/digital-experience\/wp-content\/uploads\/2026\/09\/GenUI_Protocol_Stack.png\" alt=\"Gen UI Protocol Stack\" width=\"844\" height=\"257\" srcset=\"https:\/\/blogs.infosys.com\/digital-experience\/wp-content\/uploads\/2026\/09\/GenUI_Protocol_Stack.png 936w, https:\/\/blogs.infosys.com\/digital-experience\/wp-content\/uploads\/2026\/09\/GenUI_Protocol_Stack-300x91.png 300w, https:\/\/blogs.infosys.com\/digital-experience\/wp-content\/uploads\/2026\/09\/GenUI_Protocol_Stack-768x234.png 768w\" sizes=\"auto, (max-width: 844px) 100vw, 844px\" \/><em>AG-UI is not itself a generative UI spec<\/em>\u00a0\u2014 it&#8217;s the transport layer. The actual UI description formats that ride on top of it (or work independently) are:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignleft wp-image-8852\" src=\"https:\/\/blogs.infosys.com\/digital-experience\/wp-content\/uploads\/2026\/09\/UI_Description_Formats-1.png\" alt=\"\" width=\"843\" height=\"283\" srcset=\"https:\/\/blogs.infosys.com\/digital-experience\/wp-content\/uploads\/2026\/09\/UI_Description_Formats-1.png 937w, https:\/\/blogs.infosys.com\/digital-experience\/wp-content\/uploads\/2026\/09\/UI_Description_Formats-1-300x101.png 300w, https:\/\/blogs.infosys.com\/digital-experience\/wp-content\/uploads\/2026\/09\/UI_Description_Formats-1-768x257.png 768w\" sizes=\"auto, (max-width: 843px) 100vw, 843px\" \/><\/p>\n<p><strong>A Typical Modern Pipeline<\/strong><\/p>\n<p><strong>MCP<\/strong> fetches tools\/data \u2192 the agent describes UI via <strong>A2UI<\/strong> or <strong>Open-JSON-UI<\/strong> \u2192 <strong>AG-UI<\/strong> streams that state to the frontend \u2192 a rendering library (CopilotKit, Vercel AI SDK) turns it into live components.<\/p>\n<p><strong>Generative UI Technology Landscape<\/strong><\/p>\n<ul>\n<li><strong>Frameworks<\/strong>\n<ul>\n<li>\u00a0<strong><em>Vercel AI SDK<\/em><\/strong> \u2014 AI SDK is an open-source, framework-agnostic, TypeScript toolkit that simplifies the integration of AI capabilities into modern applications. It provides powerful abstractions for handling chat interactions, streaming responses, and real-time UI updates, enabling developers to build responsive, intelligent, and engaging AI-powered user experiences with less complexity and faster development cycles.<\/li>\n<li><strong><em>CopilotKit<\/em><\/strong> \u2014 open-source framework, adopted by teams at Google, AWS, Microsoft, and LangChain. Created the AG-UI protocol. Best when you want a full &#8220;copilot&#8221; experience: chat surface, shared state, agent orchestration, and generative UI together, not just rendering.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Developer Platform<\/strong>\n<ul>\n<li><strong><em>Thesys (json-render \/ Crayon)<\/em><\/strong> \u2014 a format-authoring approach where the model outputs JSON describing a component tree. Fast to prototype with, framework-agnostic.<\/li>\n<\/ul>\n<\/li>\n<li><strong>UI Protocol<\/strong>\n<ul>\n<li><strong><em>Google A2UI<\/em><\/strong>\u00a0\u2014 focused on cross-platform, portable generative UI \u2014 not limited to a single frontend framework or the browser.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Developer Tool<\/strong>\n<ul>\n<li><strong><em>Tambo<\/em><\/strong>\u00a0\u2014 a leaner, React-only alternative to CopilotKit for teams that don&#8217;t need the full agent-orchestration surface.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Orchestration <\/strong>\n<ul>\n<li><strong><em>LangGraph \/ Mastra<\/em><\/strong>\u00a0\u2014 not UI frameworks themselves, but the most common agent-orchestration &#8220;brains&#8221; paired with the frameworks above.<\/li>\n<\/ul>\n<\/li>\n<li><strong>Standard<\/strong>\n<ul>\n<li><strong><em>MCP Apps<\/em><\/strong> \u2014 an emerging standard under the Agentic AI Foundation (Linux Foundation) for exposing generative UI through MCP tool servers. Early momentum is encouraging, but the technology has not yet reached a level of maturity that warrants broad adoption.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<p><strong>Considerations For Picking A Framework<\/strong><\/p>\n<ul>\n<li><strong>Your existing stack :<\/strong> If you\u2019re already deep in a frontend technology, a tightly integrated SDK will get you further, faster, than a framework-agnostic one. If your frontend is more heterogeneous, portability matters more than integration depth.<\/li>\n<li><strong>How much of the \u201cassistant\u201d you need, versus just the rendering :<\/strong> Some teams only need a way to render structured output as components; others want the fuller package \u2014 chat surface, agent orchestration, shared state \u2014 bundled together.<\/li>\n<li><strong>Prototype speed versus long-term flexibility :<\/strong> Formats that let the model author JSON directly tend to be the fastest way to get something on screen, but may trade off some of the control and structure you\u2019d want at scale.<\/li>\n<li><strong>How much you want the model to author versus select :<\/strong> Controlled patterns (fixed components, model just picks) are easier to reason about and secure; declarative and open-ended patterns give more expressive power at the cost of more surface area to sandbox and test.<\/li>\n<li><strong>Cross-platform needs :<\/strong> Web-only products have more options; if the same generated UI needs to show up beyond the browser, that narrows the field toward frameworks built with portability in mind.<\/li>\n<li><strong>Transport versus rendering :<\/strong> It\u2019s worth separating these two concerns \u2014 the protocol that moves agent state to the frontend doesn\u2019t have to be the same library that renders it, and decoupling them can keep you from being locked into one vendor\u2019s full stack.<\/li>\n<\/ul>\n<p>Framework selection is context-dependent. The optimal choice is determined by how much UI authoring responsibility is delegated to the model and the degree of platform independence required.<\/p>\n<p><strong>Final Thought<\/strong><\/p>\n<p>Generative UI does not eliminate frontend engineering; it changes the boundary between design-time and runtime.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Most conversations with AI still happen through a single medium: text. You ask, it [&hellip;]<\/p>\n","protected":false},"author":171,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"inline_featured_image":false,"footnotes":""},"categories":[4],"tags":[782],"coauthors":[37],"class_list":["post-8320","post","type-post","status-publish","format-standard","hentry","category-emerging-technologies","tag-generative-ui"],"acf":[],"_links":{"self":[{"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/posts\/8320","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/users\/171"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/comments?post=8320"}],"version-history":[{"count":10,"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/posts\/8320\/revisions"}],"predecessor-version":[{"id":8855,"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/posts\/8320\/revisions\/8855"}],"wp:attachment":[{"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/media?parent=8320"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/categories?post=8320"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/tags?post=8320"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/blogs.infosys.com\/digital-experience\/wp-json\/wp\/v2\/coauthors?post=8320"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}