Ronald van der Plas ยท Sep 30, 2026
Architecting for Agent Rendering
In the previous blog post, I looked at a question that started with Scrunch AXP: if AI agents are becoming another consumer of our digital experiences, should we keep transforming websites afterwards, or should we start designing our frontends with those consumers in mind from the beginning? My conclusion was fairly simple. We should not create a second website for AI agents, and we should not create a second editorial process either. The CMS should remain the single source of truth. The content stays the same, but the representation will change depending on the consumer. For a human visitor, that representation might contain navigation, carousels, tabs, personalisation, and all the other things we have spent years building into modern digital experiences. An AI agent doesn't need many of those interactions. It might benefit more from a stable, semantic and complete representation of the same information. I called the capability agent rendering. The next question is: what architecture do we need to accomplish this? We are not building two websites This is the most important constraint. If every component suddenly gets a completely separate "human version" and "agent version", editors start writing different content for agents, and every new feature needs to be implemented twice. This will surely mean the architecture has already failed. The outcome should be much simpler. The CMS provides structured content, the frontend understands that content, and the rendering layer optimises it for a particular consumer. Conceptually: CMS โ component data โ rendering layer โ representation The representation comes after the content. We are not creating another content model specifically for AI agents. Separate detection from rendering One of the first decisions is where we determine what kind of consumer is making the request. A CDN or edge layer could inspect it, middleware could do it, or the application itself could decide which representation to return. For existing enterprise websites, doing this at the edge can make a lot of sense. That is part of what makes an approach such as Scrunch AXP interesting: you can introduce agent-friendly delivery without rebuilding the systems behind it. For a new headless implementation, I would argue that we might want to take matters into our own hands. But I would not tightly couple the rendering architecture to bot detection. User agents change, new agents appear, and not every agent will identify itself in the same way. One part of the architecture should therefore determine which representation is requested, while another part knows how to render that representation. The rendering logic shouldn't care whether that decision came from a CDN, middleware, a request header or something else. The component data stays the same This is where headless architectures give us an advantage. By the time a frontend receives content from a CMS such as Sitecore or Storyblok, it already knows what that content represents. It knows that something is a hero, a product list, an accordion or a comparison component. That knowledge exists before the frontend turns the data into the visual experience we see in the browser. Take an accordion, for example: the CMS might provide five items, each containing a heading and some content. For a human visitor, we render those items as an accordion because it makes the page easier to scan. But the underlying information isn't really an accordion. It is five sections of related content. An agent renderer could simply expose those five sections as headings followed by their content. Nothing changes in the CMS or editorial process. Only the representation changes. The same applies to tabs, carousels, selectors and other interactive components. Only specialise where it is needed That doesn't mean every component needs a separate agent renderer. A heading is still a heading. A paragraph is still a paragraph. An image with useful alternative text might already be perfectly understandable in its normal HTML representation. Special handling becomes useful when human interaction makes information harder to access. An accordion is an obvious example. A carousel containing several products might be another. A comparison where information only appears after selecting different options is an even clearer case. So instead of asking how we create an AI version of every component, I think the more useful question is: Which components currently hides information behind human interaction? That gives us a much more manageable problem. What should the agent representation look like? I don't think there is a final answer yet. My starting point would still be good semantic HTML. It already works across the web and allows us to preserve hierarchy, links, media references and structured information without inventing a completely new publishing format. But the architecture shouldn't depend on semantic HTML alone. Markdown is interesting because it offers a simpler representation of textual content. Structured data can even provide extra context, and content negotiation could eventually allow a consumer to explicitly request another representation of the same resource. The important part is not whether the first version returns HTML or Markdown. It is having a clear point in the frontend where a representation can be created for a specific consumer. Keep it deterministic For the first version, I would keep the agent representation deliberately boring. No personalisation. No A/B testing. No user-specific state. No generative rewriting during the request. If the published CMS content hasn't changed, the output should ideally stay the same as well. That gives us a canonical representation which is easier to cache, test and monitor, while the approved CMS content remains the source of truth. This also prevents agent rendering from quietly becoming another content channel. The moment we start generating alternative marketing copy specifically for AI agents, we introduce new questions around approval, compliance, translations and ownership. That might become interesting later, but I wouldn't start there. Caching and testing become easier A stable representation also makes caching much more attractive. Human experiences often contain personalisation, experiments and other state that make responses vary between visitors. A public, canonical agent representation can avoid most of that. Instead of letting every agent execute the full frontend and trigger all backend calls used by the human experience, we may be able to serve a much smaller cached response. At the same time, we need to prove that simplifying the representation hasn't removed important information. Because both versions use the same source content, much of that testing can be automated. We can compare important text, links, media references, heading structure and structured data between both representations. The goal isn't to create the smallest possible page. It is to create the simplest representation that still preserves the meaning. The architecture I would propose today When designing a new headless DXP today, I wouldn't make agent rendering a completely separate subsystem. I would make it a capability of the rendering layer. The CMS would continue managing structured content, workflows, translations, permissions and publishing. The frontend would continue receiving component data, but it would no longer assume there is only one way to turn that data into an experience. Most components would have one representation (semantic HTML). Only the components where interaction hides information would need something different. Consumer detection could happen earlier at the CDN, edge or application level, while the rendering layer remains independent of how that decision was made. I am not suggesting this because we already know exactly what AI agents will need in the future. We don't. The value is that the architecture gives us a place to evolve without redesigning the frontend later. From architecture to implementation All of this still looks relatively clean on paper. The real test is what happens when we actually build it. How much additional code does an agent renderer require? Can we avoid maintaining two separate component trees? What happens to components that don't have a special agent representation? And how do we test that both versions still contain the same meaningful information? That is what I want to explore in part 3. Instead of another architecture diagram, I want to build a small agent-rendering implementation using a modern headless frontend. We can take a few typical interactive components, render them normally for a human visitor and then create a deterministic representation for an AI agent. That should show whether agent rendering remains a useful architectural idea once we actually have to maintain the code.
Read postRonald van der Plas ยท Sep 21, 2026
Your DXP was built for humans. What about AI agents?
Recently, I watched a Sitecore MVP webinar about "Meet Scrunch: The Future of AI Discovery". The session had a lot to think about, but one part in particular stuck with me: Scrunch's Agent Experience Platform, or AXP. One thing AXP can do is give AI agents a different representation of a webpage. Instead of sending all the JavaScript, tracking, interactive components and other browser-oriented complexity that we normally deliver to a human visitor, AXP can serve clean, server-rendered HTML without a JavaScript dependency. Dynamic components can also be simplified or restructured so their information is easier for an AI agent to consume. I really like that idea. Especially for large enterprise websites, where you might have years of technology decisions behind you, multiple platforms, JavaScript-heavy applications and perhaps no headless architecture yet. Then putting an intelligent delivery layer in front of what already exists would be a very pragmatic solution. You can improve what agents receive without rebuilding your entire website. But while listening to the webinar, it also made me think about the other side of this problem. If I were designing a headless DXP from scratch today, would I still want to add another platform dependency to transform my pages? Or should the frontend be able to serve AI agents properly from the start? We designed our digital experiences for humans For years, most of our assumptions have been relatively simple. Someone opens a browser, requests a page and then interacts with the experience we have carefully designed for them. You can see that assumption almost everywhere in modern websites. We use JavaScript to provide interaction. We hide information behind tabs and accordions, just to create a seamless experience. We use carousels to fit multiple products into a limited amount of space. We personalise components, run experiments, add tracking and keep making the frontend experience more sophisticated. There is nothing wrong with that. All of those things can provide real value to a human visitor. But AI agents have very different needs. It doesn't need an accordion to save screen space. It doesn't care about the animation between two slides in a carousel. It doesn't need a fancy menu to figure out where to click next in the same way a person does. It can do without all those fancy experiences that we keep adding to make our customers happy. What it does need is access to the information, plus enough structure and context to understand what that information actually means. And that is becoming more relevant. In July 2026, Cloudflare reported that automated bot traffic had surpassed human activity and was generating roughly 57% of all web requests. That includes far more than AI agents, so we shouldn't translate that number into "57% AI traffic". However, it does show that websites increasingly have machine consumers next to human ones. If that continues, our DXP architectures probably need to start acknowledging it as well. Smaller HTML is not the actual goal It is tempting to reduce this whole problem to something like "remove JavaScript for AI". I don't think that goes far enough. A small HTML document isn't automatically a useful document. Imagine an accordion containing important information, where only the first panel is available unless someone interacts with it. Removing the JavaScript does not suddenly make that page agent-friendly. It might actually make part of the content impossible to access. So the goal shouldn't simply be to create less HTML or remove more JavaScript. The goal should be to give an AI bot access to the better meaning of the content on the page. While removing all the fancy experience that the bot is not interested in. Take a product carousel with five products. A human visitor might see one product at a time and click or swipe through them. An agent doesn't need that interaction. Its representation could simply expose all five products as a semantic list. An accordion could become a series of headings and sections. Tabs could become sequential content. A complex comparison experience could expose the comparison data without requiring the interaction around it. At the same time, a normal content block with a heading and a few paragraphs probably doesn't need any special treatment at all. That is important, because I don't think the solution is to build an alternative version of every component. We only need to look at the places where the human experience gets in the way of accessing the information. Headless gives us an interesting opportunity This is where I think headless architectures become interesting. A typical implementation might have a CMS such as Sitecore or Storyblok providing structured content to a Next.js frontend. That frontend already knows what the content is, which components need to be rendered and what those components are supposed to represent. In other words, we have a lot of knowledge available before everything is turned into the human experience. An external agent-experience layer works from the other direction. It receives the finished website and transforms it back into something that is more suitable for machines. Scrunch describes AXP as sitting at the CDN layer, detecting AI traffic and serving an optimised version while leaving the normal human website unchanged. For existing websites, I think that makes a lot of sense. It allows you to retrofit agent support onto an architecture that wasn't designed for it. For a new headless architecture, however, I think we should at least ask whether some of that responsibility could live closer to the frontend. Instead of thinking only in this direction: structured content โ human experience โ machine-friendly transformation we could start thinking about: structured content โ representation appropriate for the consumer That might sound like I am proposing two websites, but that is exactly what I would try to avoid. Same content, different representation For me, this would be one of the most important architectural constraints. The CMS remains the single source of truth. I would not start by letting an AI representation invent additional content, generate new claims or maintain its own AI-specific editorial version of a page. The moment we do that, we are introducing another publishing channel, with a whole lot of maintenance nightmares. And for a large organisation, that will create a lot of questions very quickly. Which version is correct? Who approves the AI-specific text? Does it go through the same compliance process? What happens with translations? And how do we make sure the human and agent versions don't slowly start drifting apart? I would start much simpler. Same content. Same permissions. Different representation. The agent representation can reorganise existing content, expose information that was hidden behind interaction, improve its semantic structure and remove unnecessary browser machinery. But the underlying facts and access boundaries remain the same. This also makes the architecture much easier to test. We can automatically compare the human and agent representations and check whether meaningful content, links, media references or other context has disappeared. Remove the chrome, not the context Once you start looking at a page this way, there is quite a lot that an agent probably doesn't need. It is unlikely to benefit much from our account menu, cookie interface, animated navigation, tracking scripts or a footer with dozens of repetitive navigation links. Removing those things isn't only about creating cleaner HTML. It could also reduce the work performed by backend services and lower the amount of data transferred for machine traffic. But we have to be careful not to remove useful context at the same time. Breadcrumbs, for example, can tell an agent where information sits inside a site hierarchy. Brand and site identity still matter. Relevant related links can describe relationships between information. Canonical information, language information and structured data should remain available as well. I would also keep references to meaningful images. The consumer can decide whether it wants to retrieve and interpret the actual image. The same applies to video. We don't have to send an entire interactive video player just to communicate that a relevant video exists. So for me the principle is quite simple: Remove interaction and presentation noise, but preserve meaning and context. Keep the agent experience stable There is another difference between human and agent experiences that I think can work in our favour. Our human websites are increasingly dynamic. We personalise content, run A/B tests and adapt experiences based on previous behaviour or other information we have about the visitor. For the first generation of an agent representation, I would probably do almost the opposite. I would keep it boring. Make it canonical, public, deterministic and highly cacheable. No personalisation, no experimentation, no user-specific state. If the underlying content in the CMS doesn't change, the agent should ideally receive the same representation every time. That gives us something we can properly test, cache and reason about. I would also initially keep traditional search crawlers on the rendering path that we already know. Scrunch takes a similar approach with AXP and states that Googlebot and Bingbot are not routed through its agent experience, while real-time AI retrieval agents are. In the future, there might be reasons to give search engines a similar optimised representation. For existing organic search performance is simply too important to include casually in the first experiment. Don't try to fix something that isn't broken. Make it measurable This is also where I think we have to be careful with the claims we make. Removing JavaScript doesn't prove that ChatGPT will suddenly cite your website more often. Better semantic HTML doesn't automatically mean better visibility in Gemini either. Generative systems are probabilistic, and there are many factors outside our own website that influence the answers they produce. But there are things we can measure. Can we remove the client-side runtime without losing meaningful content? Can we expose information that was previously hidden behind interaction? Is the resulting structure easier for a machine to understand? Is the output stable? Is the payload smaller than the normal human version? Those are concrete engineering outcomes. From there, we can start observing what actually happens and keep improving it. Baseline, optimise, observe and repeat. For me, that is a much more useful approach than pretending we already know exactly how every AI agent wants to consume the web. Start putting this in your proposals I don't think we should wait until all standards and best practices around this have settled. If I were proposing a new headless DXP implementation today, I would already make agent readiness part of the architecture discussion. That doesn't mean promising a second rendering implementation for every component. It doesn't mean putting generative AI directly into the live request path. And it certainly doesn't mean building a second website specifically for bots. It simply means acknowledging that the browser is no longer the only consumer we have to think about. We should identify experiences where human interaction hides information. We should think about what a canonical machine representation could look like. We should keep the CMS in control of the content. And we should design the frontend in such a way that this capability can evolve when agent behaviour and web standards change. AI-assisted development is here to stay; using it makes this more interesting. Writing and maintaining an alternative representation for the small number of components that actually need one is becoming cheaper. Personally, I would much rather use AI to help developers create and test deterministic rendering code than use generative AI to rewrite approved brand content on every request. Scrunch AXP was what triggered this thought process for me, and I still think its approach makes a lot of sense, especially when you have to make an existing enterprise landscape more agent-ready without rebuilding it. But when we are designing a new headless architecture, we have the opportunity to think about this earlier. Maybe supporting both humans and machines should simply become another responsibility we expect from the frontend. I have started calling that capability agent rendering. In the next article, I will go deeper into what an agent-rendering architecture could actually look like. Where should consumer detection happen? How can components provide another representation without doubling the maintenance? How should we handle caching and testing? And where could emerging capabilities such as Markdown content negotiation fit into the picture? Because if AI agents are becoming another consumer of our digital experiences, the question is no longer only whether they can access our websites. The more interesting question is what experience we choose to give them.
Read post