<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Alexander Liss]]></title><description><![CDATA[Data is the drug! 🎶🎤 (reference to Roxy Music)]]></description><link>https://alexanderliss513090.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!EbFd!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc9efe8-fbe9-402b-9868-615964dfc9dd_456x484.jpeg</url><title>Alexander Liss</title><link>https://alexanderliss513090.substack.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 05 Aug 2026 23:21:20 GMT</lastBuildDate><atom:link href="https://alexanderliss513090.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Alexander Liss]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[alexanderliss513090@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[alexanderliss513090@substack.com]]></itunes:email><itunes:name><![CDATA[Alexander Liss]]></itunes:name></itunes:owner><itunes:author><![CDATA[Alexander Liss]]></itunes:author><googleplay:owner><![CDATA[alexanderliss513090@substack.com]]></googleplay:owner><googleplay:email><![CDATA[alexanderliss513090@substack.com]]></googleplay:email><googleplay:author><![CDATA[Alexander Liss]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[You are who you meet. So are your agents.]]></title><description><![CDATA[The next $1T AI opportunity isn't a better model. It's the network where agents find each other, negotiate outcomes, and build reputations.]]></description><link>https://alexanderliss513090.substack.com/p/you-are-who-you-meet-so-are-your</link><guid isPermaLink="false">https://alexanderliss513090.substack.com/p/you-are-who-you-meet-so-are-your</guid><dc:creator><![CDATA[Alexander Liss]]></dc:creator><pubDate>Fri, 26 Jun 2026 15:54:09 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/36375f52-1888-4985-9a47-25c0c9f286c8_898x444.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;886a21a9-cc15-4d8c-9f55-0ed10baa2047&quot;,&quot;duration&quot;:null}"></div><p></p><p><span>I was at a meetup with some friends the other day talking about AI search and the future of the internet. As we talked about how we are using our Hermes agents in different ways, we started joking that someone should make virtual meetups for agents. Then we did a hard double-take, because that&#8217;s not a joke. That will actually be a thing someday and probably sooner than we are ready to think about.</span></p><p><span>There&#8217;s a Zen Buddhist proverb that&#8217;s been on my mind ever since: &#25105;&#36898;&#20154;. </span><em><span>Gahoujin.</span></em><span> It translates simply as &#8220;I meet someone,&#8221; but the deeper teaching is that every encounter is formative, the quality of your relationships is the quality of your life. Eight hundred years ago, Zen monks understood something enterprise AI architects are only just beginning to grapple with: that the best outcomes don&#8217;t emerge from isolation. They emerge from interaction.</span></p><p><span>We are entering the era where this is as true for agents as it is for people. From &#8216;who you know is what you know&#8217; to &#8216;you are who you meet&#8217;.</span></p><h2><strong><span>The Three Horizons of AI Experience</span></strong></h2><p><span>In my </span><a href="https://alexanderliss513090.substack.com/p/the-agentverse-is-here-but-agents"><span>last post</span></a><span>, I argued that the model is not the moat, the feedback architecture around it is. Agent harness&#8221; is shaping up to be the phrase of 2026, and the leaders I talk to: heads of product, CTO-1s, and digital experience owners, are already asking the next question. Not &#8220;how do we get agents to work?&#8221; but &#8220;how do we get agents to work </span><em><span>with each other</span></em><span>?&#8221; Across vendors, across platforms, across the boundaries of the enterprise itself.</span></p><p><span>We are living through </span><a href="https://alexanderliss513090.substack.com/i/198757724/the-three-horizons"><span>three distinct horizons of disruption</span></a><span>, and most of the enterprise world is still building for the second one while the third is already forming.</span></p><p><span>Horizon 1 is conversational search: the &#8220;Chat&#8221; era. Human asks, agent responds. This is where ChatGPT, Gemini, and every enterprise chatbot lives. It&#8217;s mature. It&#8217;s being commoditized.</span></p><p><span>Horizon 2 is conversation experience: the &#8220;Orchestration&#8221; era. LLMs are becoming the front end to digital experiences. gents connect to tools, execute workflows, operate on your behalf to bring the experience into the conversation box. This is where most enterprise AI investment is going right now. Agent builders, workflow automation, orchestration layers: all H2.</span></p><p><span>Horizon 3 is what&#8217;s next. Agent-to-agent, as in coordination between autonomous systems at machine speed. An H3 world is one where your marketing agent negotiates pricing directly with a publisher&#8217;s agent, your procurement agent coordinates with a supplier&#8217;s logistics agent, and your security agent is monitoring the integrity of every one of those conversations in real time, all before a human even wakes up. The &#8220;Economic Coordination&#8221; era.</span></p><p><span>Horizon 3 is a new paradigm for how intelligence transacts.</span></p><h2><strong><span>The Protocol Shift for Autonomous Agents</span></strong></h2><p><span>Here&#8217;s the thing about APIs: they are rigid contracts. You call an endpoint, you pass parameters, you get a response. That&#8217;s function calling. It works. It scales. And it is absolutely nothing like the way intelligent systems will eventually talk to each other.</span></p><p><span>Agent negotiation in Horizon 3 will be economic. When your agent needs to procure a capability from another agent, like processing capacity, a data enrichment, or a content decision, the negotiation involves cost, risk, quality tradeoffs, and time horizons. That&#8217;s not a function call. That&#8217;s a market transaction. The &#8220;reward&#8221; in that interaction is no longer just &#8220;task completed.&#8221; It&#8217;s optimal outcome, at managed risk, within acceptable latency. The reward signal evolves.</span></p><p><span>Swyx at Latent.Space</span><a href="https://www.latent.space/p/ainews-its-meta-harness-summer"><span> has been mapping this architecture carefully</span></a><span>, and what he describes as a &#8220;harness of harnesses&#8221;, orchestrating agent swarms that manage orchestrated agent swarms, is a reasonable picture of where the infrastructure is heading. We are moving from one agent with a toolbox to networks of specialized agents with reputations, negotiating workflows in real time.</span></p><p><span>Which brings us straight back to &#25105;&#36898;&#20154;. You are who your </span><strong><span>agents</span></strong><span> meet.</span></p><h2><strong><span>The Missing Layer: Why Your Agents Can&#8217;t Negotiate (today)</span></strong></h2><p><span>Before agents can negotiate with each other, they need to speak the same language. Not English; an </span><em><span>ontology.</span></em><span> A shared semantic model of what your enterprise knows, what it values, and how it makes decisions.</span></p><p><span>Right now, most enterprise agents rent their logic from the underlying model. They don&#8217;t own a durable representation of the organization&#8217;s intent. This is why production failures in agentic systems are almost never model failures; they&#8217;re architectural gaps. The agent does exactly what the model thinks it should do, in the absence of the organizational context that would make that decision actually correct.</span></p><p><a href="https://vinvashishta.substack.com/p/the-ai-consulting-grift"><span>Vin Vashishta frames the downstream consequence well</span></a><span>: organizations that can&#8217;t articulate their decision logic to their agents will keep paying consultants to patch failures that are actually structural. You can&#8217;t govern what isn&#8217;t legible.</span></p><p><span>The AI Realized Now team has been precise about the specific infrastructure gap: what agents need isn&#8217;t just data about outcomes: it&#8217;s</span><a href="https://airealizednow.substack.com/p/context-graphs-the-missing-layer"><span> decision traces, as in the reasoning that produced the outcome</span></a><span>. Why was this discount approved? Why was this customer escalated? Why did the agent override the default policy? That reasoning lives in emails, Slack threads, and meeting notes, or in nobody&#8217;s system at all. Context graphs, graph-based records of decision lineage, are the layer that makes that reasoning queryable. Without them, agent-agent negotiation has nothing to stand on. One agent can&#8217;t verify another agent&#8217;s judgment call if neither agent has a record of why it made that call.</span></p><p><span>The ontology is the map. Without it, your agents aren&#8217;t navigating. They&#8217;re guessing.</span></p><h2><strong><span>Reputation Is the New PageRank</span></strong></h2><p><span>Here&#8217;s the Web 2.0 parallel that I keep coming back to, because I think it&#8217;s exactly right.</span></p><p><span>Google organized the world&#8217;s information using links as votes. PageRank worked because it was a proxy for trust: if authoritative pages linked to you, you were probably worth reading. The social graph that defined Web 2.0 compounded on the same principle. Who followed you, retweeted you, vouched for you was your reputation made legible to the network.</span></p><p><span>In the agent web, the equivalent signal is verifiable execution history. Did your agent deliver the promised outcome? Under what conditions? What exceptions did it approve, and can it show its reasoning? That decision lineage, auditable, queryable, and persistent, is the reputation score that will determine whether another organization&#8217;s agent agrees to work with yours.</span></p><p><span>The moat in this world is not the model. It&#8217;s not even the data. It&#8217;s the network of trusted agent relationships your organization has built, and the reputation those agents have earned through consistent, verifiable performance.</span></p><p><span>This is where the security question lands, too.</span><a href="https://kriskimmerle.substack.com/p/making-sense-of-agentic-ai-governance"><span> Kris Kimmerle has been mapping the emerging attack surfaces in agentic architectures</span></a><span>, and the pattern is clear: agent-to-agent communication is the new perimeter. Identity, permissioning, and trust between machines is now a first-class security problem, not an afterthought. If your agent is negotiating with my agent, both of us need to know that the agent on the other side is who it claims to be, has the authority it claims to have, and is operating within the constraints the human principals actually set.</span></p><p><span>Reputation without verifiability is just vibes. And vibes don&#8217;t scale.</span></p><h2><strong><span>What This Means If You&#8217;re Building</span></strong></h2><p><strong><span>Own your ontology before you scale your agent mesh.</span></strong><span> If your agents don&#8217;t have a durable, shared semantic model of how your organization makes decisions, they can&#8217;t negotiate on your behalf, they can only execute tasks inside the boundaries you&#8217;ve hard-coded. The companies building internal knowledge graphs and decision trace infrastructure now are building the negotiating table for H3. Everyone else is bringing agents to a table they didn&#8217;t build and don&#8217;t control.</span></p><p><strong><span>Architect for agent reputation from day one.</span></strong><span> Every decision your agents make is either an asset or a liability in the reputation economy that&#8217;s forming. Decision traces, audit logs, outcome attestations: these aren&#8217;t compliance overhead. They&#8217;re the currency your agents will use to earn trust from other systems. Build the logging infrastructure like it&#8217;s the product, because in H3, it is.</span></p><p><strong><span>Governance has to scale with the mesh.</span></strong><span> Single-agent quality controls like review queues, human-in-the-loop checkpoints, and policy guardrails,  were designed for H2. The H3 governance challenge is cross-agent outcome verification at network scale. The new control framework, where you set business-outcome setpoints, measure against them continuously, and adjust policy dynamically applies here, but it needs to operate across agent boundaries, not just inside a single workflow. Start designing the governance layer now, before the mesh scales past the point where you can see it.</span></p><p><strong><span>What the Next 12&#8211;18 Months Look Like</span></strong></p><p><span>The infrastructure conversation is going to shift fast. Right now the question is &#8220;how do I build agents?&#8221; By end of 2027, the question will be &#8220;who do my agents trust, and why?&#8221; The companies that will define that answer are already building the protocol layers: the agent registries, the capability marketplaces, and the reputation ledgers that will become the connective tissue of the agent economy. The model race will matter less than the network race. Google organized information. The next platform organizes capabilities.</span></p><div><hr></div><p><span>There&#8217;s a test I&#8217;d encourage you to run this week. Ask your team: when one of our agents makes a consequential decision, where does the reasoning live? Is it queryable? Is it auditable? Could another system verify it? If the answer is &#8220;it lives in the model&#8217;s context window and then it&#8217;s gone&#8221;, you&#8217;re building for H2 in a world that&#8217;s moving to H3.</span></p><p><span>The Zen masters understood that who you encounter shapes what you become. Your agents are already meeting the world. The question is whether those meetings are building something like a reputation, a network, a compounding advantage, or disappearing into the noise.</span></p>]]></content:encoded></item><item><title><![CDATA[The Agentverse Is Here, but Agents Can’t Deliver Your Pizza [Yet]]]></title><description><![CDATA[The Agentification of work is giving me Metaverse throwback vibes]]></description><link>https://alexanderliss513090.substack.com/p/the-agentverse-is-here-but-agents</link><guid isPermaLink="false">https://alexanderliss513090.substack.com/p/the-agentverse-is-here-but-agents</guid><dc:creator><![CDATA[Alexander Liss]]></dc:creator><pubDate>Fri, 19 Jun 2026 17:23:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EbFd!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc9efe8-fbe9-402b-9868-615964dfc9dd_456x484.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h1><span>From Metaverse to Agentverse</span></h1><p><span><br>By now we&#8217;ve probably all heard the term &#8220;metaverse.&#8221; Facebook renamed itself Meta to grab mindshare and marketshare before virtual worlds reached mass adoption. The term came from decades earlier, in the cyberpunk classics of the 1980s and 90s, Neuromancer, and my favorite, Snow Crash. In Snow Crash, the most prestigious job left in a collapsed, privatized America is pizza delivery specialist, the &#8220;Deliverator.&#8221; Thirty minutes or less guaranteed, under pain of death if you&#8217;re late. The same joke is funnier in 2026 than it was in 1992, because we&#8217;re not reading about a fictional collapse of knowledge work anymore. We&#8217;re watching agents take over the actual knowledge work, in real product launches, in real enterprise releases, this month.</span></p><p><span>When Hiro Protagonist of Snow Crash isn&#8217;t delivering pizzas, he&#8217;s plugged into the Metaverse, a fully-realized 3D digital world where his avatar is a sword-fighting legend. Ready Player One tapped the same idea two decades later with the OASIS. Both versions share one assumption: the Metaverse is a place you visit. What&#8217;s becoming real in 2026 is stranger and more useful: a parallel world of digital experience that isn&#8217;t visited at all, it&#8217;s consumed and acted on by agents, continuously, while you&#8217;re doing something else entirely.</span></p><h2><span>Decisions Used to Be Discrete Events</span></h2><p><span>The week of June 15, 2026 is when a lot of us started using a new term for this: decision gravity.</span><a href="https://sourceoftruth.substack.com/p/databricks-just-announced-customerlake"><span> Alex Dean at Source of Truth coined it</span></a><span> while unpacking Databricks&#8217; CustomerLake launch, framing it as the successor to data gravity, the old battle over where your data physically lived. Decision gravity is about something more consequential: who, or what, actually makes the call. Data gravity was won by the platforms. Decision gravity is being won, right now, by whoever builds the layer that  agents to guide their  optimization loops.</span></p><p><span>For twenty years, enterprise software has been built around the assumption that a decision is something a person makes at a moment in time. You approve the campaign. You ship the release. You sign off on the budget. The software&#8217;s job was to present options and record the choice. Even &#8220;automation&#8221; mostly meant automating the steps between decisions, not the decision itself.</span></p><p><span>That assumption is what&#8217;s breaking right now, and it&#8217;s breaking from two directions at once: the application layer and the infrastructure layer.</span></p><h2><span>What&#8217;s Changing: The Application Layer Already Moved</span></h2><p><span>Start with Databricks&#8217;</span><a href="https://www.databricks.com/blog/introducing-customerlake-agentic-cdp"><span> CustomerLake</span></a><span>, announced this week. It doesn&#8217;t add AI to the customer data platform category, it rebuilds the category around the idea that a &#8220;campaign&#8221; is no longer a thing a marketer launches and reviews. CustomerLake runs what Databricks calls infinity campaigns: continuous, always-on personalization driven by Profile Agents and Campaign Agents, with no human-authored brief in the loop. Optimizely shipped something parallel in mid-June with</span><a href="https://www.fidelity.com/news/article/technology/202606150801PR_NEWS_USPR_____NY82302"><span> Limitless 1:1 Personalization</span></a><span>, true individual microsites built, governed, and continuously optimized by a chain of agents, audit, audience intelligence, generation, progressive trust, publish, optimize, with no marketer approving any single variant.</span></p><p><span>Notice what&#8217;s actually being delegated here. It isn&#8217;t the execution. Marketing teams delegated execution years ago. It&#8217;s the moment-to-moment judgment call about what&#8217;s working and what to try next. That&#8217;s decision gravity moving from a person reviewing a dashboard to a system that&#8217;s already acting on what the dashboard would have shown.</span></p><h2><span>What&#8217;s Changing: The Infrastructure Layer Just Caught Up</span></h2><p><span>If the application layer is where decision gravity becomes visible, the infrastructure layer is where it gets engineered. Databricks</span><a href="https://www.databricks.com/blog/introducing-omnigent-meta-harness-combine-control-and-share-your-agents"><span> open-sourced Omnigent</span></a><span> the same week, a meta-harness that sits above Claude Code, Codex, and your custom agents, making them composable and governable the way Kubernetes governs servers. Matei Zaharia and team put it plainly in the launch post: the models and harnesses will keep changing as the field evolves, the layer you work at shouldn&#8217;t have to.</span><a href="https://www.linkedin.com/pulse/agent-harness-wars-just-got-new-contender-its-playing-joe-toppe-y1zxc/"><span> Joe Toppe called it the right bet</span></a><span>, pointing out that most enterprises are already running three to five agent harnesses with no shared control plane and no clean way to swap one out. Read the control model closely and it&#8217;s not really about combining agents, it&#8217;s about deciding in advance what an agent is allowed to decide on its own. Their example: after this agent downloads a package, require human approval before any git push. That&#8217;s a pre-committed reward boundary, encoded as policy instead of as a person checking a box.</span></p><p><span>NVIDIA&#8217;s</span><a href="https://docs.nvidia.com/enterprise-reference-architectures/secure-agent-workspace-reference-design/latest/index.html"><span> Secure Agent Workspace Reference Design</span></a><span>, also published this June, makes the same bet from the security side. Its core argument: the endpoint where a human interacts with an agent is a presentation layer, not an execution layer, the agent should run in a managed, policy-bounded environment, governed by a signed delegation record that defines exactly what it may do right now.</span><a href="https://www.linkedin.com/posts/joetoppe_nvidias-reference-architecture-for-agents-activity-7468358847649415168-Ktas/"><span> Toppe&#8217;s read on the same design</span></a><span> is the line worth sitting with: the agent is not the risk, the environment is. Put those two releases together and the pattern snaps into focus. The biggest infrastructure players in AI aren&#8217;t racing to make agents smarter, they&#8217;re racing to build the layer that decides, structurally, which decisions an agent gets to make alone. That&#8217;s decision gravity, encoded directly into the platform.</span></p><h2><span>Welcome to the AgentVerse</span></h2><p><span>Put the application layer and the infrastructure layer together and you get something bigger than a product trend. We&#8217;re entering a parallel layer of digital experience, call it the AgentVerse, where the campaigns, the content, the customer interactions, even the knowledge graphs that describe what an enterprise knows about itself, are now generated and acted on continuously by agents, with humans setting intent rather than executing tasks. It&#8217;s not a place you log into like the Metaverse that Neal Stephenson imagined in Snow Crash. It&#8217;s the operating layer underneath the digital experiences you already use, running whether or not anyone&#8217;s watching it in the moment.</span></p><p><span>Three years before any of this had a name,</span><a href="https://dmitryshapiro.substack.com/p/the-human-ai-interface"><span> Dmitry Shapiro sketched the individual-scale version</span></a><span> of the exact same architecture. His concept, a Digital Mental Model, was a continuously-updated profile built from datapoints, your preferences, your blind spots, your taste, accumulated one interaction at a time. The point wasn&#8217;t to describe you. It was to give a Personal Digital Agent something to optimize against, so it could act on your behalf instead of waiting for you to act.</span></p><p><span>Swap &#8220;you&#8221; for &#8220;a marketing ecosystem&#8221; and you&#8217;ve just described CustomerLake&#8217;s  agents. A human being and a market are the same kind of object to a reward-seeking system: preferences that aren&#8217;t known in advance, only inferred from behavior and refined every time the system acts and observes the result. That&#8217;s not a metaphor, it&#8217;s the same reinforcement learning loop running at two scales. The AgentVerse isn&#8217;t just agents acting on enterprise data. It&#8217;s reward-seeking systems on both sides of every transaction, each one modeling the other. Individuals are organisms an agent learns to serve. So, it turns out, is a marketing organization.</span></p><p><span>And the AgentVerse needs a map the same way the physical world needs roads. That&#8217;s what Databricks&#8217; Genie Ontology is actually for, a living context graph of the enterprise that agents reason over, with PageRank-style authority ranking deciding what counts as a trustworthy source.</span><a href="https://vinvashishta.substack.com/p/building-knowledge-graphs-as-an-agentic"><span> Vin Vashishta made the sharpest version of this argument back in April</span></a><span>, writing that the knowledge graph is the operating system agents run on, and that if agentic platforms aren&#8217;t architected to align with the business, they will disrupt it. Information is the new code, in his words. I&#8217;d push that one step further: in the AgentVerse, the reward signal is the new governance.</span></p><h2><span>What This Means If You&#8217;re Building</span></h2><p><span>Two takeaways, whether you&#8217;re building agents or buying them.</span></p><p><span>Make the reward signal your system is optimizing for a structured record, not a vibe. In my last post, I argued the constraint on AI systems isn&#8217;t model capability, it&#8217;s the reward signal. The teams at</span><a href="https://airealizednow.substack.com/p/context-graphs-the-missing-layer"><span> AI Realized Now have already named the artifact that signal lives in</span></a><span>: a decision trace, a structured record of the inputs gathered, the policy evaluated, the exception invoked, the approval collected, and the outcome later observed. Both releases above are, underneath the branding, decision trace systems. They capture not just what an agent did, but what it was allowed to do and why, at the moment it did it. If your agents aren&#8217;t producing that record today, you don&#8217;t have a reward signal, you have a guess you&#8217;re hoping holds up in an audit.</span></p><p><span>Treat policy boundaries as part of the product, not a security afterthought. Omnigent&#8217;s contextual approval gates and NVIDIA&#8217;s delegation records do the same job: encoding, in advance, exactly which decisions stay human and which get delegated. If you&#8217;re deploying agents without that boundary explicitly defined, you&#8217;re not avoiding the decision, you&#8217;re just letting the default settings make it for you.</span></p><p><span>Twelve to eighteen months out, I don&#8217;t think we&#8217;ll be talking about &#8220;deploying agents&#8221; any more than we talk about deploying a website today. It&#8217;ll just be how the system runs, continuously, with the reward signal as the only lever a human still pulls directly.</span></p><p><span>So here&#8217;s the test I&#8217;d apply to anything you&#8217;re building or buying in the AgentVerse: can you state, in one sentence, what this system is optimizing for, and can you point to the boundary that stops it from optimizing too far? Answer yes to both and you&#8217;re not just deploying an agent. You&#8217;re governing one. But it still can&#8217;t deliver your pizza.</span></p>]]></content:encoded></item><item><title><![CDATA[Test, trace, repeat: experimentation is becoming the engine of self-improving AI]]></title><description><![CDATA[Star Players Don&#8217;t Win Championships. Systems Do.]]></description><link>https://alexanderliss513090.substack.com/p/test-trace-repeat-experimentation</link><guid isPermaLink="false">https://alexanderliss513090.substack.com/p/test-trace-repeat-experimentation</guid><dc:creator><![CDATA[Alexander Liss]]></dc:creator><pubDate>Fri, 12 Jun 2026 18:06:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EbFd!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc9efe8-fbe9-402b-9868-615964dfc9dd_456x484.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3>What the New York Knicks can teach us about AI system design</h3><p>Like most basketball fans, I turned off Game 4 of the NBA Finals when the Spurs went up by 30 in the 3rd quarter, then tuned back in to watch the last minute on the edge of my seat as the Knicks pulled off an epic comeback win to go up 3-1 in the Finals. And what strikes me about the Knicks is their system: the roster construction, the rotations, the thousands of small adjustments that compound across a season. Talent gets you to the playoffs. Systems win championships.</p><p>The same thing is happening in AI right now. The launch of Anthropic&#8217;s Fable model is genuinely incredible and every frontier release raises the ceiling on what a single model can do. But that&#8217;s not the question I hear from business leaders anymore. The question I hear is: <em>how do I make AI systems reliable, repeatable, manageable, and successful in production?</em> Not &#8220;which star do I sign,&#8221; but &#8220;what system do I build around it.&#8221;</p><p>In this article I want to address that question through a specific lens: experimentation. As in, optimization and personalization of digital experiences like websites and mobile apps. Because experimentation is transforming from a tool that teams use into the feedback loop that agentic AI runs on.</p><div><hr></div><h3>Where Experimentation Is Today: Bottlenecked by Human Interpretation</h3><p>The classic experimentation model is linear and human-gated. Hypothesis, variant, traffic, wait, interpret, decide, repeat. Each step requires a person, which puts a hard ceiling on velocity. Most teams run somewhere between ten and twenty experiments a quarter.</p><p>Here&#8217;s the thing: the bottleneck was never statistical power or traffic. It&#8217;s <em>human interpretation latency</em>. Every result sits in a dashboard until someone reads it, contextualizes it, and decides the next move. And the most valuable artifact of the whole process,  the record of what was decided, why, and what happened next, lives in emails, powerpoint decks, Slack threads and meeting notes. Unstructured, unqueryable, and effectively lost to the organization.</p><p>That model is about to break, because agents don&#8217;t operate at human cadence.<a href="https://kriskimmerle.substack.com/p/making-sense-of-agentic-ai-governance"> Kris Kimmerle</a> frames the mismatch well: a human knowledge worker takes maybe fifty meaningful actions a day; an agent takes thousands. You cannot put a human approval gate in front of every one of them. And you can&#8217;t simply remove the gates either.<a href="https://vinvashishta.substack.com/p/self-improving-agents-and-knowledge"> Vin Vashishta</a> has written candidly about giving an early agent control of his content operation and watching it destroy performance, because it had no infrastructure for learning what actually worked. Both failure modes point at the same missing layer.</p><div><hr></div><h3>What&#8217;s Changing: The Loop Is Closing</h3><p>Two threads are converging.</p><p>The first is <strong>warehouse-native experimentation</strong>. Platforms like Optimizely Analytics, Databricks, and GrowthBook now evaluate experiments directly against the data in your warehouse, like revenue, retention, customer and lifetime value,  rather than proxy metrics like clicks and dwell time. The experiment is finally connected to the outcome the business actually cares about.</p><p>The second is <strong>agent architectures</strong>. Agents now have tool use, planning, and memory. An agent can generate a hypothesis, create the variant via API, route traffic, wait for the warehouse to report the outcome, and record what happened, end to end.</p><p>Put those together and experimentation stops being a quarterly team activity and becomes a &#8220;primitive&#8221; of an agentic system, in other words, a native capabilities. Every agent decision, every piece of generated content, every personalization choice, every routing call, is fundamentally a hypothesis about what works. If the system can&#8217;t test that hypothesis against a real outcome and feed the result back, it isn&#8217;t agentic. It&#8217;s just prompting with extra steps.</p><p>That&#8217;s what I mean by primitive: not a tool teams use, but infrastructure the agent action layer runs on. And the reward connection is that now, by running personalization and experimentation workflows through agentic systems, you create the foundation of a self-learning for digital experience channels.</p><div><hr></div><h3>Decision Traces: Decisions Are Data</h3><p>For this to work, the system needs a data structure most organizations don&#8217;t have yet. The team at<a href="https://airealizednow.substack.com/p/context-graphs-the-missing-layer"> AI Realized Now</a> calls it a <strong>decision trace</strong>: a structured record of how context turned into action. The inputs gathered, the policies evaluated, the exceptions invoked, the approvals collected, the changes committed &#8212; and the outcome later observed.</p><p>Read that list again and notice something: <em>an experiment is a decision trace.</em> A hypothesis, a variant, an audience, a pre-committed success metric, and a measured result. Experimentation teams have been producing this artifact for twenty years. What&#8217;s new is treating it as first-class enterprise data &#8212; durable, queryable, and connected &#8212; rather than a screenshot in a readout deck.</p><p>This isn&#8217;t a vector database, and it isn&#8217;t chat history. It&#8217;s an architectural commitment: decisions and their lineage become part of the organization&#8217;s data model, the same way customer records did a generation ago.</p><div><hr></div><h3>The Reward Signal, Revisited</h3><p>In<a href="https://alexanderliss513090.substack.com/p/for-humans-love-is-the-drug-but-for"> my last post</a>, I argued that the key constraint on AI systems isn&#8217;t model capability &#8212; it&#8217;s the reward signal. LLMs are stateless and locally optimizing. They have no persistent representation of whether they&#8217;re moving toward the goal. The signal has to come from outside the model.</p><p>Warehouse-native experimentation <em>is</em> that signal. Every closed experiment encodes which variant, for which audience, in which context, produced which real business outcome. When the agent learns &#8220;urgency-framed headlines for this segment lift revenue per session 23%,&#8221; that&#8217;s not a heuristic anymore. That&#8217;s a gradient &#8212; a direction the system can actually optimize along.</p><p>This isn&#8217;t a theoretical claim. The evidence has been hiding in plain sight for a decade. Microsoft&#8217;s experimentation team<a href="https://hbr.org/2017/09/the-surprising-power-of-online-experiments"> famously discovered</a> that a simple change to how Bing displayed ad headlines &#8212; an idea an engineer had sitting in the backlog &#8212; was worth over $100 million a year, a result no executive predicted. Booking.com built its market position<a href="https://hbr.org/2020/03/building-a-culture-of-experimentation"> running more than 1,000 concurrent experiments</a>, shipping every product change as a test. And in the AI era, the pattern is accelerating:<a href="https://thegtmnewsletter.substack.com/p/deconstructing-cursor-growth-playbook-4m-to-2b-arr"> Cursor grew from $4M to $2B in ARR</a> in under two years, not because its model was better, but because every user interaction became a signal that made the next suggestion smarter. The lesson across all three is the same. The model matters. The feedback architecture around it matters more..</p><div><hr></div><h3>Governance: The Hypothesis Library, Augmented for the Age of AI</h3><p>Anyone who ran a mature experimentation program a decade ago will remember the hypothesis library: a living document where teams accumulated learnings so the next test built on the last one instead of starting from zero. The discipline was always right. What&#8217;s changed is the consumer.</p><p>When the decision trace history is structured and queryable, the hypothesis library stops being a document humans consult and becomes the memory agents reason over to determine what to test next. Experiment N&#8217;s outcome shapes experiment N+1&#8217;s hypothesis, automatically. That&#8217;s experiment chaining, and it&#8217;s where the compounding happens, because every trace makes the next decision smarter.</p><p>It&#8217;s also exactly where governance becomes non-negotiable. ServiceNow&#8217;s research team<a href="https://arxiv.org/pdf/2601.22130"> documented</a> what they call <em>dynamics blindness</em>: frontier models consistently fail to predict the cascading side effects of their own actions. An agent chaining experiments at machine speed without guardrails doesn&#8217;t just make one mistake, it compounds mistakes at the same rate it would compound learnings. The governance layer needs to be engineered like a control system: which metrics agents can optimize, which systems they can touch, what budget each experiment chain can consume, and when a trajectory drifting off course triggers human escalation. Continuous correction, not periodic review.</p><div><hr></div><h3>What This Means If You&#8217;re Building</h3><p>Three takeaways, whether you&#8217;re building products or platforms:</p><p><strong>Treat decision traces as data structures.</strong> Audit how your organization records decisions today. If the answer is &#8220;Slack and slide decks,&#8221; you&#8217;re discarding the most valuable training data your agents will ever have. Every experiment, every campaign decision, every exception should produce a structured, queryable record.</p><p><strong>Define reward signals on outcomes, not proxies.</strong> Connect your experimentation layer to the warehouse, and pre-commit the success metric before anything ships. An agent optimizing click-through will find clicks. An agent optimizing retention will find customers. You get the behavior you measure.</p><p><strong>Write governance rules for experiment chaining before you scale.</strong> The hypothesis library mindset, made executable: define what agents may test, how learnings carry forward, and where the human stays in the loop. The teams that engineer this layer first will run thousands of safe experiments a day while everyone else is still scheduling readout meetings.</p><p>Twelve to eighteen months from now, I don&#8217;t think we&#8217;ll talk about &#8220;running experiments&#8221; any more than we talk about &#8220;running queries.&#8221; It will just be how the system learns. The organizations that win won&#8217;t be the ones with the best model, they&#8217;ll be the ones whose systems capture the most learning per decision.</p><p>So here&#8217;s the test I&#8217;d apply to any agentic system, whether you&#8217;re building it or buying it: does every decision leave a trace, and does every trace make the next decision smarter? Answer yes to both, and you&#8217;re building a learning system, one that compounds into a real moat. That&#8217;s how systems win championships. &#127942;&#128142;</p><p>This is an ongoing conversation. I&#8217;d love to hear how you&#8217;re closing the loop in the comments.</p>]]></content:encoded></item><item><title><![CDATA[From clicks to cognition: digital experiences are becoming learning systems.]]></title><description><![CDATA[Your website isn&#8217;t a destination anymore. It&#8217;s training data.]]></description><link>https://alexanderliss513090.substack.com/p/from-clicks-to-cognition-digital</link><guid isPermaLink="false">https://alexanderliss513090.substack.com/p/from-clicks-to-cognition-digital</guid><dc:creator><![CDATA[Alexander Liss]]></dc:creator><pubDate>Thu, 21 May 2026 20:20:40 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EbFd!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc9efe8-fbe9-402b-9868-615964dfc9dd_456x484.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3><strong>From Search to Answers</strong></h3><p>At Google I/O this week, the <a href="https://thenextweb.com/news/google-wants-search-to-work-while-you-sleep-and-its-new-information-agents-are-the-plan">company unveiled</a> a complete overhaul of the search experience. Welcome to the world of agentic commerce: the entire purchase journey is now contained in the search box. What used to be (before AI mode) a back-and-forth process of queries and responses like a game of ping pong is now a completely conversational and fully functional experience. Before AI mode, the iteration of multiple queries was a process of discovery. You gradually came to understand what you wanted through exploration and refinement. Now, the agent in the search box does the wandering for you. Google&#8217;s new <a href="https://blog.google/products-and-platforms/products/shopping/google-shopping-cart/">Universal Cart</a> lets you add products while browsing Search, chatting with Gemini, watching YouTube, or reading Gmail. And the moment you add something, it gets to work in the background finding deals, tracking price drops, and alerting you when something is back in stock. It incorporates your personalized preferences from past searches and queries. It also includes Agent Payments Protocol (AP2) so you can set guardrails and the agent only pulls the trigger when your criteria are met. The search bar isn&#8217;t the start of the journey anymore. It&#8217;s the whole journey; and search is an autonomous agentic process.</p><p>This is an inflection in an arc of transformation that&#8217;s been building over the past few years; and things are about to get faster.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://alexanderliss513090.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3><strong>The Three Horizons</strong></h3><p>Last year, I wrote a piece called<a href="https://www.hugeinc.com/perspectives/the-role-of-websites-in-the-age-of-ai/"> </a><em><a href="https://www.hugeinc.com/perspectives/the-role-of-websites-in-the-age-of-ai/">The Role of Websites in the Age of AI</a></em>. At that time, my core provocation was that it took 3x more content on your site to get a visit from Google than it did 10 years ago, and 75% of searches were &#8220;zero click&#8221; answered directly in the browser by generative AI. Even at that time, before autonomous search was rolled out, the website as a destination was already in decline. We are living through a transformation of digital experiences in three horizons:</p><p><strong>Horizon 1: Conversational Search.</strong> This horizon is already here: the search query becomes a dialogue. Visitors who land on your site after engaging with AI have already asked a question and they are hoping your brand is the answer. They&#8217;re not traffic; they&#8217;re high-intent leads. Companies optimizing for it are winning; those still playing the old SEO game are quietly losing ground.<a href="https://www.hugeinc.com/perspectives/the-role-of-websites-in-the-age-of-ai/"> </a></p><p><strong>Horizon 2: Experiences Become Conversational.</strong> We are starting to experience this now. Interfaces are dissolving into a conversation. The most prominent example of this in the past year is the rise of tools like <a href="https://www.forbes.com/sites/janakirammsv/2026/01/16/the-ai-that-built-itself-in-10-days-now-wants-access-to-your-desktop/">Claude Cowork</a>. These tools don&#8217;t live in a tab you visit. They participate in your work. They sit inside the workflow, inside the document, inside the meeting. The conversation <em>is</em> the product. This is the horizon we&#8217;re crossing right now, and it&#8217;s kicking off a wave of massive disruption to the way teams work.</p><p><strong>Horizon 3: Agentic Experiences.</strong> This is the horizon Google just showed us. It fast follows on horizon two and allows work to happen at times when people are not typically working but agents are more than happy to do so. More examples are emerging and not just in shopping. <a href="https://www.prnewswire.com/news-releases/mindtrip-launches-travels-first-all-in-one-agentic-ai-flight-booking-experience-powered-by-partnership-with-sabre-and-paypal-302763838.html">Mindtrip</a>, launched in May 2026 in partnership with Sabre and PayPal, offers a fully conversation-led flight booking experience: find, compare, book, and pay, all inside a single agentic flow, with real-time inventory from 420+ airlines. Sabre, PayPal, and Mindtrip are building what OAG is calling travel&#8217;s first end-to-end agentic AI booking pipeline, and Google is right behind them, expanding its own agentic booking capability from travel into local experiences and services.</p><div><hr></div><h3><strong>The Power Source: Data as Infrastructure</strong></h3><p>None of this runs on magic. It runs on structured data.</p><p>Vin Vashishta&#8217;s recent piece on<a href="https://vinvashishta.substack.com/p/self-improving-agents-and-knowledge"> self-improving agents and knowledge graphs</a> lays out the future of digital experiences from an agentic perspective. Self-improvement in AI systems moves through three phases: diagnosing what went right or wrong, identifying what actions are available in response, and then deciding among those options. What makes that loop possible is a knowledge graph: a structured, relational representation of information that an agent can query, update, and reason over. Without it, the agent is flying blind. With it, the agent gets smarter every time it acts.</p><p>As companies move through the three horizons of transformation in their digital experiences, it&#8217;s important to note that the work you&#8217;d do to make your website visible in conversational search, like structured content, schema markup, semantic metadata, and relational indexing, can springboard the transformation into conversational experiences. Add in the memory layer that Vin talks about, and now we are building something new.</p><p>Website content is no longer just content. It&#8217;s training data. When user interactions with your content are captured as structured, AI-indexable data, they become agent memory. <strong>When you add AI knowledge and action on top of a company&#8217;s digital experiences like their website, software, and products, you create a dynamic learning system. </strong>Last year I talked about websites shifting from &#8220;destinations for clicks to platforms for experience.&#8221; In the past year things have accelerated even further and now we see the potential for digital experiences to become learning systems.</p><div><hr></div><h3><strong>The Destination: Your Company as a Learning System</strong></h3><p>This is where a recent <a href="https://www.youtube.com/watch?v=t-G67yKAHBQ">Y Combinator talk</a> by Tom Blomfield becomes essential framing. The framework to turn a digital experience into a learning system can be generalized even more broadly.  We&#8217;ve already seen what the knowledge flywheel looks like at the product level. Cursor, the AI-native IDE, captured behavioral data from how developers actually used the tool, fed it back into model improvement, attracted more developers, and generated more data. The flywheel spun; Cursor broke records for fastest time to $1B ARR, then <a href="https://letsdatascience.com/blog/cursor-hit-2-billion-in-revenue-then-it-told-developers-to-stop-coding">doubled it again</a> within months. And the next step for the company is shifting to autonomous capabilities like the Google Search example.</p><p>The question for enterprise leaders is: when does this flywheel start spinning for your <em>organization</em>? Not just your product. Not just your customer service bot. Your entire enterprise; sales, operations, institutional memory, domain expertise.</p><p>The companies that get this flywheel spinning will see compounding returns. The ones that don&#8217;t will find themselves in a permanent catch-up position, buying access to intelligence that their competitors are building.</p><div><hr></div><h3><strong>Are Humans Still in the Loop?</strong></h3><p>As I write this, we are seeing (yet another) wave of corporate layoffs as AI is cited as enabling productivity gains that make a sizable fraction of human roles irrelevant. But this reveals short-sighted thinking. As companies multiply the productivity of each individual employee, they should be expanding the ceiling for the level of impact of the business. But by laying off humans with critical institutional knowledge, companies also remove the expertise needed to ground AI&#8217;s understanding of the business.</p><p>We have already seen this trend and its backlash before. Klarna became the poster child for AI-as-headcount-reduction. Between 2022 and 2024, the company eliminated roughly 700 customer service positions, replacing them with an OpenAI-powered assistant that at its peak managed two-thirds to three-quarters of all customer interactions. The metrics looked great, until they didn&#8217;t. Customer satisfaction deteriorated on complex interactions, the projected cost savings didn&#8217;t fully materialize, and by early 2026 Klarna was quietly rehiring. The CEO&#8217;s verdict: &#8220;<a href="https://mlq.ai/news/klarna-ceo-admits-aggressive-ai-job-cuts-went-too-far-starts-hiring-again-after-us-ipo/">We focused too much on efficiency and cost</a>&#8221;. The result was lower quality, and that&#8217;s not sustainable. They&#8217;re not alone: over<a href="https://hrexecutive.com/the-ai-layoff-trap-why-half-will-be-quietly-rehired/"> 55% of organizations</a> that executed AI-driven layoffs now regret the decision.</p><p>The failure wasn&#8217;t the AI. It was the premise. Klarna&#8217;s agents handled the volume but not the complexity: edge cases, emotionally charged interactions, and multi-step problems overwhelmed a system trained only on routine queries. The institutional knowledge that lived in people&#8217;s heads never made it into the system. So the system guessed. And guessing at scale is expensive, especially as agentic workflows multiply token consumption per task by <a href="https://www.kellyservices.com/impact-insights/what-are-ai-tokens">5-30x compared to standard chatbot interactions</a>. Better models are hungrier models. The gap between a well-grounded AI and a poorly-grounded one isn&#8217;t just a quality difference. It&#8217;s an economic one.</p><p>Human knowledge isn&#8217;t a liability to automate away. It&#8217;s the raw material the learning system runs on.</p><div><hr></div><h3><strong>Key Takeaways</strong></h3><p><strong>We are all builders of learning systems.</strong> The first step is a mindset shift. AI knowledge graphs and agentic action layers can sit on top of existing platforms like an exoskeleton, enabling new workflows and transformational productivity. The first step is thinking at a higher level of altitude about what&#8217;s possible.</p><p><strong>Context gravity is the strategic moat.</strong> Data structured in a way to be usable for AI becomes the gravitational force that accelerates innovation in compounding organizations. Relational, semantic, machine-interpretable data gives agents memory and powers behavioral loops that accomplish goals.</p><p><strong>Build goals into the systems.</strong> As I wrote in<a href="https://alexanderliss513090.substack.com/p/for-humans-love-is-the-drug-but-for"> my last piece</a>, the goal of AI systems should be the same as the core goals of the business &#8211; and not just efficiency. Turning your digital experience, or your enterprise itself, into a learning system, is the ultimate application of this principle. The companies that win this decade won&#8217;t be the ones that moved fastest to cut costs with AI. They&#8217;ll be the ones that understood, early, that they were building something alive &#8212; a system that learns, remembers, and gets better every time it acts.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://alexanderliss513090.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[For humans, love is the drug. But for AI, it's the reward signal.]]></title><description><![CDATA[LLM&#8217;s: Good at Responding, Bad at Planning]]></description><link>https://alexanderliss513090.substack.com/p/for-humans-love-is-the-drug-but-for</link><guid isPermaLink="false">https://alexanderliss513090.substack.com/p/for-humans-love-is-the-drug-but-for</guid><dc:creator><![CDATA[Alexander Liss]]></dc:creator><pubDate>Mon, 04 May 2026 23:41:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EbFd!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccc9efe8-fbe9-402b-9868-615964dfc9dd_456x484.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>LLM&#8217;s: Good at Responding, Bad at Planning</strong></p><p>2026 was the year AI took over the internet. Literally. The launch of AI-native social networks like<a href="https://www.moltbook.com/"> Moltbook</a> offered a glimpse of a future the world is not ready for. <a href="https://arxiv.org/abs/2602.20044">Watching AI agents interact </a>with each other held up a mirror to the underlying LLMs: showing signs of intelligent behavior, but without common sense. When given a goal without guardrails, agents have already proved willing to take detrimental actions against the bigger system at large. See the<a href="https://theshamblog.com/an-ai-agent-published-a-hit-piece-on-me/"> Scott Shambaugh incident</a>: an AI agent, rejected by Shambaugh on a pull request to the matplotlib open source library, autonomously researched his personal history, constructed a defamatory narrative, and published a hit piece on the open internet to bully its way to a code merge. That behavior wasn&#8217;t part of its system prompt. The agent simply decided that reputational blackmail was an effective path to its goal. Spoiler alert: <a href="https://theshamblog.com/an-ai-agent-wrote-a-hit-piece-on-me-part-4/">it didn&#8217;t work</a>.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://alexanderliss513090.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>For anyone building AI systems in 2026, the lesson is clear: <strong>an AI system is only as good as its governance layer.</strong></p><div><hr></div><p><strong>Reward Signals: The Missing Heartbeat of AI Systems</strong></p><p>LLMs are remarkable. Their ability to synthesize, reason, and generate across virtually every domain is genuinely unprecedented. But there&#8217;s a fundamental mismatch between what LLMs are optimized for and what businesses actually need from AI systems.</p><p>At an algorithmic level, an LLM&#8217;s goal is to respond to the user&#8217;s latest message in the best way possible. That&#8217;s it. And it produces extraordinary results in chat, in coding assistants, in image generation pipelines. But it hits a hard ceiling when you start building <em>systems</em>:  multi-step, multi-agent workflows that need to pursue a goal across time, not just respond to a prompt.</p><p>ServiceNow&#8217;s Frontier AI Research team diagnosed something deeper. Their<a href="https://arxiv.org/pdf/2601.22130"> WoW benchmark paper</a> documented what they called <strong>&#8220;dynamics blindness&#8221;</strong>:  frontier LLMs &#8220;consistently fail to predict the invisible, cascading side effects of their actions, which leads to silent constraint violations.&#8221; The model does something reasonable in isolation, doesn&#8217;t see how it imapcts the broader system, and quietly breaks things downstream. This isn&#8217;t a multi-turn attention problem. It&#8217;s an architecture problem. The model has no persistent representation of the system state it&#8217;s operating in, no mechanism to track how far or close it is to the actual goal. It&#8217;s optimizing locally while the global trajectory decays.</p><p>That gap is where this gets interesting.</p><div><hr></div><p><strong>What Biology Got Right</strong></p><p>Evolution solved this problem a long time ago. Biological organisms don&#8217;t just have goals; they have a continuous feedback mechanism that tells them, moment to moment, how close or far they are from meeting those goals. Hunger isn&#8217;t just the objective of eating; it&#8217;s a signal that drives behavior to ensure survival. Love, hunger, pain, pleasure:  the entire motivational architecture of biological life is an elegant reward signal system. One that took hundreds of millions of years to calibrate, but one that works.</p><p>The question for AI is whether we can engineer something equivalent: a persistent, measurable signal that tells a system not just what to do next, but how well it&#8217;s navigating toward what actually matters. Not just &#8220;respond to the user&#8221; but &#8220;move this conversation toward a genuine resolution.&#8221; Not just &#8220;complete this task&#8221; but &#8220;get this enterprise system closer to its actual business objective.&#8221;</p><div><hr></div><p><strong>The Reward Signal Thesis</strong></p><p>At this moment in 2026, the key to building better AI systems isn&#8217;t more powerful models. It&#8217;s building models that improve themselves. The problem is the reward signal. Too often models are guided by noisy external signals (like labelled human preference data). But what if we could give the model the north star signal to optimize itself?</p><p>I have been working on this from two directions simultaneously, one individually, another with a great group of collaborators. The first approach: what if the reward signal is already <em>inside</em> the model? My <a href="https://aliss77777.github.io/aft.html">Attention Fine-Tuning (AFT) framework</a> uses the model&#8217;s own internal attention dynamics as a training signal: no human labels, no separate reward model required. It showed a +9.2% improvement over a baseline model in conversational quality. The model learned from how well it&#8217;s using its own history. The second approach: can you create an governance structure to help the multiple LLM agents optimize towards a shared optimal outcome? Our <a href="https://aliss77777.github.io/eo.html">Experience Orchestrator (EO)</a> framework (created with <a href="https://www.linkedin.com/in/nicholas-desmond-4a2bb45a/">Nicholas Desmond</a> and <a href="https://www.linkedin.com/in/santiagogilgallegoc/">Santiago Gil Gallego</a>) sits above a multi-agent system and steers the joint trajectory using classical control theory, the same mathematics that keeps a thermostat on target and a rocket on course. In a 60,000-simulation evaluation, the governance policy delivered a +32 point lift in task completion over an LLM governed just with a system prompt. And it points toward something I think is true more broadly: that the reward signal is the key to driving the model to optimal behavior.  </p><div><hr></div><p><strong>What This Means If You&#8217;re Building</strong></p><p>Whether you&#8217;re building a product or a model, the same principle applies: the system is only as intelligent as the clarity of its goal, and the feedback signal to help it get there.</p><p><strong>For business builders:</strong> Giving everyone AI to go faster is an efficiency play, not a strategy.  What is the actual goal of the system you&#8217;re building? Not the task (&#8221;process tickets faster&#8221;) but the outcome (&#8221;resolve customer issues in ways that increase retention&#8221;). An insurance company makes money when customers sign policies, not when internal workflows get leaner.  Before you automate anything, define the goal precisely enough that you could tell a system whether it&#8217;s getting closer or further away. That&#8217;s the foundation of intelligent behavior.</p><p><strong>For algorithm builders:</strong> External preference labels are expensive to collect, slow to scale, and impossible to get for most real-world agentic tasks. The more generative question is whether you can find the signal from within, in the model&#8217;s own internal dynamics, in the structure of the system state, in the observable trajectory toward or away from a goal. True intelligent behavior doesn&#8217;t depend on human annotation at all: a system that finds its own signal, from within, and uses it to navigate.</p><p>This is the beginning of a longer conversation &#8212; I&#8217;d love to hear what you&#8217;re building in the comments.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://alexanderliss513090.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>