<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Product Theatre]]></title><description><![CDATA[Honest writing on AI, product leadership, and the gap between hype and reality — one pattern, every Tuesday.]]></description><link>https://www.producttheatre.com</link><image><url>https://substackcdn.com/image/fetch/$s_!kIaA!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4b557c19-2676-4227-9c7c-37092ec14f3f_512x512.png</url><title>Product Theatre</title><link>https://www.producttheatre.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 21 Jul 2026 07:30:52 GMT</lastBuildDate><atom:link href="https://www.producttheatre.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Cam]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[producttheatre@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[producttheatre@substack.com]]></itunes:email><itunes:name><![CDATA[Cam]]></itunes:name></itunes:owner><itunes:author><![CDATA[Cam]]></itunes:author><googleplay:owner><![CDATA[producttheatre@substack.com]]></googleplay:owner><googleplay:email><![CDATA[producttheatre@substack.com]]></googleplay:email><googleplay:author><![CDATA[Cam]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Free Isn't Free]]></title><description><![CDATA[In AI, every free user costs you money. Price for that.]]></description><link>https://www.producttheatre.com/p/free-isnt-free</link><guid isPermaLink="false">https://www.producttheatre.com/p/free-isnt-free</guid><dc:creator><![CDATA[Cam]]></dc:creator><pubDate>Mon, 29 Jun 2026 21:02:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!IZex!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IZex!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IZex!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg 424w, https://substackcdn.com/image/fetch/$s_!IZex!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg 848w, https://substackcdn.com/image/fetch/$s_!IZex!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!IZex!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IZex!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg" width="900" height="507" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:507,&quot;width&quot;:900,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A pricing ladder where the lowest rung glows hot, with a melting GPU casting a long shadow.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A pricing ladder where the lowest rung glows hot, with a melting GPU casting a long shadow." title="A pricing ladder where the lowest rung glows hot, with a melting GPU casting a long shadow." srcset="https://substackcdn.com/image/fetch/$s_!IZex!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg 424w, https://substackcdn.com/image/fetch/$s_!IZex!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg 848w, https://substackcdn.com/image/fetch/$s_!IZex!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!IZex!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7be513d5-5f49-4d6f-bb5d-eb45ad14f377_900x507.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In SaaS, an extra free user costs you nothing. In AI, every free user costs you money.</p><p>Every time they hit Enter, your GPUs fire and your cash burns. Growth explodes. Then your bills arrive.</p><p>So the old freemium playbook breaks. &#8220;Give away the basics, gate the best features&#8221; assumes free is cheap. In AI, your free tier is your single biggest compute bill, and your best feature might be the most expensive thing you own.</p><p>A huge thank-you to Vikas Kansal and the Google AI team, whose paywall framework this is built on. &#128591;</p><h2>The trap: one premium tier</h2><p>Google AI hit this wall in public. Their first move was the classic play: a single $20 Gemini Advanced tier, pay for the smartest model.</p><p>Two things broke. The free tier was already, in users&#8217; own words, &#8220;smarter than I am,&#8221; so most saw no reason to upgrade. And the power users who did upgrade burned so much compute the unit economics were terrifying.</p><p>In other words: one tier can&#8217;t price a cost that scales with every prompt. You have to gate on what users value and what the company pays for, at the same time.</p><p>So what do you gate? Three things.</p><h2>Gate the volume</h2><p>Price the work pumped through the system, not the smarts. Google split one tier into three: Plus, Pro and Ultra, each a level of usage intensity, up to a 1M-token context window. Midjourney does the same with Fast Mode (instant GPU, metered hours) versus Relax Mode (free, but you queue). Light users stay free. Heavy users pay because they cost more, not just because they value more.</p><h2>Gate the outcome</h2><p>Stop selling answers. Start selling hours.</p><p>The free tier gives the right answer, then leaves you to copy, paste, reformat and re-prompt. Put the paywall in front of the features that finish the job: automation, agents, integrations, export. Intercom&#8217;s Fin agent is the cleanest version. It&#8217;s free to let the AI try, and you pay $0.99 only when the problem is actually resolved.</p><h2>Gate the heaviest compute</h2><p>Some features melt the servers. When Google built Genie 3, its real-time world model, the internal joke was that &#8220;the TPUs were melting on every prompt.&#8221; Serving that to every free user wasn&#8217;t a bad business move. It was physically impossible.</p><p>So it went to the top tier only. Make text and basic images universally free to pull people in. Set a hard gate the moment someone wants cinematic video, a real-time simulation, or a 3D world.</p><h2>Tiers capture the budget. The ecosystem keeps it.</h2><p>Right now, users are pouring experimental dollars into AI. Those budgets won&#8217;t last. Well-priced tiers capture that money. An ecosystem around them is what keeps it.</p><p><strong>Convert at the moment of intent.</strong> The upsell is timing, not packaging. Google watches three signals: a user who refines the same output five times in one session (that&#8217;s real work), a user on both desktop and mobile inside 48 hours (the tool is now part of their day), and the &#8220;continue this chat&#8221; soft paywall, where a shared conversation needs a Pro model and proves its value before anyone pays.</p><p><strong>Bundle for month two.</strong> AI churn is brutal because the habit isn&#8217;t formed yet. So tie the subscription to something stickier. Google bundled AI with Google One cloud storage, and nobody cancels the thing holding their photos, so the AI habit survives almost by accident. Cursor did it by indexing your codebase: churning means tearing down your own setup.</p><p><strong>Route prompts so cheap stays cheap.</strong> You can&#8217;t serve your biggest model to every free prompt. &#8220;What&#8217;s the capital of France?&#8221; should hit a tiny, fast model. A logic puzzle routes to a heavy reasoner with token metering. The user still gets instant magic. Your margin stays invisible.</p><h2>The traps underneath</h2><p>Three margin traps will catch you even when the gates are right.</p><p><strong>Don&#8217;t break trust.</strong> AI usage spikes around projects and exams, then stops. Bury the cancel button and a temporary pause becomes a permanent exit. Offer a one-click pause that keeps their saved prompts, and they come back for the next heavy week.</p><p><strong>Don&#8217;t lock your tiers in stone.</strong> Today&#8217;s premium model is next quarter&#8217;s commodity. Lock the boundaries and you bleed margin (giving away $5,000 of compute on a $200 plan) or bleed users (a rival gives your &#8220;premium&#8221; away free). Audit unit costs constantly. Keep room at the top for the next breakthrough.</p><p><strong>Don&#8217;t ship peak-agnostic pricing.</strong> A flat 24/7 rate while your GPUs redline on weekday afternoons and sit idle on Sundays trains people to run heavy jobs at your most expensive hour. Add usage multipliers or compute &#8220;happy hours&#8221; to smooth the load.</p><p>Gate the volume. Gate the outcome. Gate the heaviest compute.</p><p>Open your pricing page next to your compute bill. If your most expensive feature sits in your cheapest tier, you&#8217;ve found the leak. Tell me which feature it is. I read every reply. &#128591;</p><h2>Source Notes</h2><ul><li><p><strong>AI freemium cost structure.</strong> Vikas Kansal&#8217;s Google AI essay via Lenny&#8217;s Newsletter supplies the central claim: SaaS freemium assumes low marginal cost, while AI freemium has real inference cost on every active user. </p></li><li><p><strong>Volume, outcome, and compute gates.</strong> The Google AI examples, Gemini tiering, and Genie 3 compute constraint come from Kansal&#8217;s pricing framework. Source: </p></li></ul><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:195672025,&quot;url&quot;:&quot;https://www.lennysnewsletter.com/p/why-saas-freemium-playbooks-dont&quot;,&quot;publication_id&quot;:10845,&quot;publication_name&quot;:&quot;Lenny's Newsletter&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!8MSN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441213db-4824-4e48-9d28-a3a18952cbfc_592x592.png&quot;,&quot;title&quot;:&quot;Why SaaS freemium playbooks don&#8217;t work in AI, and what to do instead&quot;,&quot;truncated_body_text&quot;:&quot;&#128075; Hey there, I&#8217;m Lenny. Each week, I answer reader questions about building product, driving growth, and accelerating your career. For more: Lenny&#8217;s Podcast | Lennybot | How I AI | My favorite AI/PM courses, public speaking course, and interview prep copilot&quot;,&quot;date&quot;:&quot;2026-05-05T13:03:32.007Z&quot;,&quot;like_count&quot;:346,&quot;comment_count&quot;:6,&quot;bylines&quot;:[{&quot;id&quot;:131847289,&quot;name&quot;:&quot;Vikas Kansal&quot;,&quot;handle&quot;:&quot;vikaskansal&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e59fdf91-22ef-4b6b-9b84-35e86d7c9641_400x400.jpeg&quot;,&quot;bio&quot;:&quot;Product lead for Google AI subscriptions, leading a team of PMs building AI monetization for consumers. AI subscriptions include the best of Google AI, including Gemini App, NotebookLM and cloud storage!&quot;,&quot;profile_set_up_at&quot;:&quot;2026-04-27T21:02:25.552Z&quot;,&quot;reader_installed_at&quot;:null,&quot;is_guest&quot;:true,&quot;bestseller_tier&quot;:null,&quot;status&quot;:{&quot;bestsellerTier&quot;:null,&quot;subscriberTier&quot;:null,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:null,&quot;subscriber&quot;:null},&quot;primaryPublicationId&quot;:8927213,&quot;primaryPublicationName&quot;:&quot;Vikas Kansal&quot;,&quot;primaryPublicationUrl&quot;:&quot;https://vikaskansal.substack.com&quot;,&quot;primaryPublicationSubscribeUrl&quot;:&quot;https://vikaskansal.substack.com/subscribe?&quot;}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.lennysnewsletter.com/p/why-saas-freemium-playbooks-dont?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!8MSN!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441213db-4824-4e48-9d28-a3a18952cbfc_592x592.png" loading="lazy"><span class="embedded-post-publication-name">Lenny's Newsletter</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">Why SaaS freemium playbooks don&#8217;t work in AI, and what to do instead</div></div><div class="embedded-post-body">&#128075; Hey there, I&#8217;m Lenny. Each week, I answer reader questions about building product, driving growth, and accelerating your career. For more: Lenny&#8217;s Podcast | Lennybot | How I AI | My favorite AI/PM courses, public speaking course, and interview prep copilot&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">3 months ago &#183; 346 likes &#183; 6 comments &#183; Vikas Kansal</div></a></div><ul><li><p><strong>Aha moment without unlimited compute.</strong> Elena Verna&#8217;s freemium analysis supports the point that the free tier still has to deliver enough value for activation, even when cost must be controlled. Source: </p></li></ul><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:181378441,&quot;url&quot;:&quot;https://www.elenaverna.com/p/why-ai-doesnt-mean-the-end-of-freemium&quot;,&quot;publication_id&quot;:1435249,&quot;publication_name&quot;:&quot;Elena's Growth Scoop&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!Ex2M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19a83a66-b8ab-4490-b665-0ecc789b8947_96x96.png&quot;,&quot;title&quot;:&quot;Why AI doesn&#8217;t mean the end of Freemium&quot;,&quot;truncated_body_text&quot;:&quot;I&#8217;ve been a Freemium fan forever. It&#8217;s the best. But LLMs are expensive, so does that mean freemium is dead? A lot of products seem to think so - locking every AI feature behind a paywall like it&#8217;s 2014 SaaS all over again.&quot;,&quot;date&quot;:&quot;2025-12-12T17:14:45.170Z&quot;,&quot;like_count&quot;:57,&quot;comment_count&quot;:5,&quot;bylines&quot;:[{&quot;id&quot;:3478323,&quot;name&quot;:&quot;Elena Verna&quot;,&quot;handle&quot;:&quot;plgrowth&quot;,&quot;previous_name&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6eeab89f-c508-46ec-91a1-9e8e3ce3021a_1926x2892.jpeg&quot;,&quot;bio&quot;:&quot;Always be learning. \nCurrently doing growth at dropbox. Previously surveymonkey, miro, amplitude, netlify. Advised mongodb, clockwise, sanity.io, krisp, and many more. &quot;,&quot;profile_set_up_at&quot;:&quot;2021-11-10T19:00:26.905Z&quot;,&quot;reader_installed_at&quot;:&quot;2023-02-24T20:44:47.139Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:1398663,&quot;user_id&quot;:3478323,&quot;publication_id&quot;:1435249,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:1435249,&quot;name&quot;:&quot;Elena's Growth Scoop&quot;,&quot;subdomain&quot;:&quot;elenaverna&quot;,&quot;custom_domain&quot;:&quot;www.elenaverna.com&quot;,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;A must-have resource for anyone in the growth, marketing, or product space. &quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/19a83a66-b8ab-4490-b665-0ecc789b8947_96x96.png&quot;,&quot;author_id&quot;:3478323,&quot;primary_user_id&quot;:3478323,&quot;theme_var_background_pop&quot;:&quot;#121BFA&quot;,&quot;created_at&quot;:&quot;2023-02-20T21:19:16.215Z&quot;,&quot;email_from_name&quot;:&quot;Elena's Growth Scoop&quot;,&quot;copyright&quot;:&quot;Elena Verna&quot;,&quot;founding_plan_name&quot;:&quot;'I Can Expense This' &quot;,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;enabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}}],&quot;twitter_screen_name&quot;:&quot;ElenaVerna&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:100,&quot;status&quot;:{&quot;bestsellerTier&quot;:100,&quot;subscriberTier&quot;:null,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;bestseller&quot;,&quot;tier&quot;:100},&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://www.elenaverna.com/p/why-ai-doesnt-mean-the-end-of-freemium?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!Ex2M!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19a83a66-b8ab-4490-b665-0ecc789b8947_96x96.png" loading="lazy"><span class="embedded-post-publication-name">Elena's Growth Scoop</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">Why AI doesn&#8217;t mean the end of Freemium</div></div><div class="embedded-post-body">I&#8217;ve been a Freemium fan forever. It&#8217;s the best. But LLMs are expensive, so does that mean freemium is dead? A lot of products seem to think so - locking every AI feature behind a paywall like it&#8217;s 2014 SaaS all over again&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">7 months ago &#183; 57 likes &#183; 5 comments &#183; Elena Verna</div></a></div><ul><li><p><strong>Outcome pricing.</strong> Intercom&#8217;s Fin pricing is the clean example of charging when the AI resolves the customer problem, not merely when it generates an answer. Source: <a href="https://www.intercom.com/help/en/articles/8205718-fin-ai-agent-outcomes">https://www.intercom.com/help/en/articles/8205718-fin-ai-agent-outcomes</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Model Is Not the Bottleneck]]></title><description><![CDATA[AI value is stuck in the gap between tool adoption and changed work.]]></description><link>https://www.producttheatre.com/p/the-model-is-not-the-bottleneck</link><guid isPermaLink="false">https://www.producttheatre.com/p/the-model-is-not-the-bottleneck</guid><dc:creator><![CDATA[Cam]]></dc:creator><pubDate>Wed, 17 Jun 2026 04:48:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YVYn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YVYn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YVYn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg 424w, https://substackcdn.com/image/fetch/$s_!YVYn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg 848w, https://substackcdn.com/image/fetch/$s_!YVYn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!YVYn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YVYn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg" width="900" height="507" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:507,&quot;width&quot;:900,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Two clocks on a board: the model clock racing ahead while the work clock moves through workflows, incentives, controls, and trust.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two clocks on a board: the model clock racing ahead while the work clock moves through workflows, incentives, controls, and trust." title="Two clocks on a board: the model clock racing ahead while the work clock moves through workflows, incentives, controls, and trust." srcset="https://substackcdn.com/image/fetch/$s_!YVYn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg 424w, https://substackcdn.com/image/fetch/$s_!YVYn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg 848w, https://substackcdn.com/image/fetch/$s_!YVYn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!YVYn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f5454af-9dd8-48e3-87fb-53ed3829f01a_900x507.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Someone in your company is still asking the wrong AI question.</p><p>&#8220;Which model should we use?&#8221;</p><p>&#8220;Which agent platform should we buy?&#8221;</p><p>&#8220;How many people are using it?&#8221;</p><p>Those questions are not useless. They are just not the hard part anymore.</p><p>The hard question is:</p><p><strong>What work will change because of this?</strong></p><p>That is the integration gap: the space between AI being available and AI changing how work happens.</p><p>Adoption means people used the tool. Integration means the work changed because of it.</p><p>Most companies are counting adoption and hoping it proves transformation.</p><p>It does not.</p><h2>The model clock is faster than the work clock</h2><p>There are two clocks running inside every company now.</p><p>The model clock moves in releases, benchmarks, demos, context windows, coding gains, and product drops.</p><p>It is fast. It is public. It is easy to see.</p><p>The work clock moves through incentives, leaders, workflows, data access, controls, trust, job design, quality standards, and one political question most AI roadmaps avoid:</p><p>Who has to change how they work?</p><p>That clock is slower.</p><p>Sometimes it barely moves.</p><p>When the model clock keeps accelerating and the work clock stays stuck, companies do not get transformation.</p><p>They get more impressive activity.</p><h2>The last mile was the whole road</h2><p>The clearest signal is not another model release.</p><p>It is the model companies moving into deployment.</p><p>In 2026, OpenAI launched the OpenAI Deployment Company to help businesses build around intelligence. Anthropic announced an enterprise AI services company with Blackstone, Hellman &amp; Friedman, and Goldman Sachs.</p><p>Strip away the launch language and the admission is plain:</p><p>Access is not enough.</p><p>If access were enough, frontier labs would sell the interface and wait. Instead, they are moving toward workflows, data, use cases, handoffs, controls, permissions, and change management.</p><p>That is where the value is stuck.</p><p>Not in the model.</p><p>In the work around the model.</p><p>The last mile was not adoption.</p><p>It was integration.</p><h2>Adoption is not integration</h2><p>Adoption is easy to count.</p><p>Seats assigned. Training completed. Pilots launched. Champions nominated. Usage up. Internal posts saying, &#8220;I made this with AI.&#8221;</p><p>That can be a good start.</p><p>But it is not proof that anything important changed.</p><p>A person can use AI every day and still leave the organisation untouched. They can draft faster, summarise faster, code faster, analyse faster, and still push the work through the same old system.</p><p>The same meeting.</p><p>The same handoff.</p><p>The same approval path.</p><p>The same quality bar.</p><p>This is the difference:</p><p><strong>Adoption</strong> &#8594; <strong>Integration</strong></p><ul><li><p>People use the tool &#8594; The workflow changes</p></li><li><p>Usage goes up &#8594; A task, handoff, or decision changes</p></li><li><p>Champions share tips &#8594; Teams reset the baseline</p></li><li><p>Pilots prove possibility &#8594; Controls make scale possible</p></li><li><p>Leaders ask for examples &#8594; Leaders change the operating model</p></li></ul><p>Adoption is contact with the tool.</p><p>Integration is change in the work.</p><h2>The power-user story is not the strategy</h2><p>The most dangerous companies are not the ones doing nothing.</p><p>They are the ones where smart individuals have moved faster than the system around them.</p><p>You can see the pattern.</p><p>A product leader has a private research workflow. An engineering leader has a code-review loop. A finance analyst has a forecasting prompt chain.</p><p>Each one is real.</p><p>Each one saves time.</p><p>Each one makes the person look sharp.</p><p>Then it dies there.</p><p>No one turns it into a reusable workflow. No one defines the quality bar. No one wires it to governed data. No one asks whether this should become the default way the team works.</p><p>The company gets AI anecdotes instead of AI infrastructure. That is the real failure mode: not lack of enthusiasm, but lack of integration.</p><h2>The leader layer decides what sticks</h2><p>If you want one diagnostic, look at leaders.</p><p>AI becomes real work when leaders change what good work looks like.</p><p>A leader who uses AI visibly gives permission. A leader who checks AI-assisted work sets the quality bar. A leader who only asks for usage numbers teaches the team to perform adoption.</p><p>This is why leader-on-tools matters.</p><p>Not because every executive needs to become a prompt engineer.</p><p>Because a leader who has never sat with the tools cannot set the standard for the work the tools produce.</p><p>They cannot tell the difference between a saved hour and a redesigned workflow, a clever demo and a production system, a prompt trick and a new team baseline.</p><p>The leader layer is where adoption either becomes the work or stays a side project.</p><h2>Change the dashboard</h2><p>Stop asking only whether people used AI.</p><p>Ask whether the work changed.</p><p>Use one integration review with five questions:</p><ol><li><p><strong>What workflow changed?</strong> &#8212; Whether AI removed, redesigned, or improved a real step in the work</p></li><li><p><strong>What baseline moved?</strong> &#8212; Whether one person's better method became normal for the team</p></li><li><p><strong>What quality bar changed?</strong> &#8212; Whether leaders can judge the output beyond "the tool produced it"</p></li><li><p><strong>What controls are in place?</strong> &#8212; Whether data access, evaluation, logging, policy, and escalation can support scale</p></li><li><p><strong>What decision or outcome improved?</strong> &#8212; Whether the change made work faster, better, cheaper, safer, or more reliable</p></li></ol><p>These questions are industry-agnostic. They work for product surfaces, internal workflows, service operations, marketplace processes, regulated decisions, and engineering systems.</p><p>If the answer is only &#8220;people saved time,&#8221; the value story is still soft.</p><p>Saved time matters.</p><p>But the company has to decide what that saved time becomes.</p><h2>Replace the adoption update with an integration review</h2><p>Pick one AI initiative on the roadmap.</p><p>Replace the adoption update with an integration review.</p><p>Do not ask:</p><p>&#8220;How many people used it?&#8221;</p><p>Ask:</p><p><strong>What work changed, who owns the new baseline, and what proof shows a better decision or outcome?</strong></p><p>If the team cannot answer, the initiative is not transformation yet.</p><p>It is adoption.</p><p>That is not a failure.</p><p>It is just not the finish line.</p><p>The models are moving.</p><p>The question is whether the institution can move with them.</p><p>Do not buy a faster model and call it strategy.</p><p>Change the work.</p><h2>Source Notes</h2><ul><li><p><strong>Adoption versus integration.</strong> PwC&#8217;s AI performance work supports the claim that value concentrates in companies that change the operating model around AI, not only adopt tools. External source: <a href="https://www.pwc.com/gx/en/issues/technology/ai-performance/want-ai-roi-go-for-growth.html">https://www.pwc.com/gx/en/issues/technology/ai-performance/want-ai-roi-go-for-growth.html</a></p></li><li><p><strong>Leader and organisation effects.</strong> Microsoft&#8217;s Work Trend Index supports the point that AI impact depends on agents, human agency, and organisational conditions, not only individual usage. External source: <a href="https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization">https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization</a></p></li><li><p><strong>Deployment-company signal.</strong> OpenAI and Anthropic moving into enterprise deployment services are used as evidence that access alone is no longer the hard part. Sources: <a href="https://openai.com/index/openai-launches-the-deployment-company/">https://openai.com/index/openai-launches-the-deployment-company/</a> and <a href="https://www.anthropic.com/news/enterprise-ai-services-company?pubDate=20260225">https://www.anthropic.com/news/enterprise-ai-services-company?pubDate=20260225</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Benchmark Fallacy]]></title><description><![CDATA[AI beats us on every test. That was never the question.]]></description><link>https://www.producttheatre.com/p/the-benchmark-fallacy</link><guid isPermaLink="false">https://www.producttheatre.com/p/the-benchmark-fallacy</guid><dc:creator><![CDATA[Cam]]></dc:creator><pubDate>Thu, 11 Jun 2026 01:15:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KZLK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KZLK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KZLK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg 424w, https://substackcdn.com/image/fetch/$s_!KZLK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg 848w, https://substackcdn.com/image/fetch/$s_!KZLK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!KZLK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KZLK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg" width="900" height="507" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:507,&quot;width&quot;:900,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A teenager holding a perfect driving-test score in one hand and a set of ambulance keys in the other, looking unsure.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A teenager holding a perfect driving-test score in one hand and a set of ambulance keys in the other, looking unsure." title="A teenager holding a perfect driving-test score in one hand and a set of ambulance keys in the other, looking unsure." srcset="https://substackcdn.com/image/fetch/$s_!KZLK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg 424w, https://substackcdn.com/image/fetch/$s_!KZLK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg 848w, https://substackcdn.com/image/fetch/$s_!KZLK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!KZLK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F512a4203-49a8-4461-91fd-8f38aafc0204_900x507.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>AI can now pass almost any test we give it. So people say it has caught up to us.</p><p>The sharpest warning about that came from an odd place. Not a tech lab. The Vatican.</p><p>Pope Leo XIV wrote a letter that captures it exactly. A machine can beat you on every test you can write, and still not be able to replace you.</p><p>The trap has a name worth keeping. The <strong>benchmark fallacy</strong>: treating &#8220;AI beat the human on the test&#8221; as &#8220;AI can replace the human.&#8221; (A benchmark is just a scored test.)</p><p>A test grades the answer. It never grades what makes an answer good when things go sideways: judgment, and the nerve to own the result. Passing the test and holding the job are two different claims. We keep merging them.</p><h2>Acing the test is not the same as being trusted with the job.</h2><p>Picture a sixteen-year-old who aces the written driving test. The score is real. It proves they learned the rules.</p><p>It does not mean you would hand them the keys to an ambulance in a snowstorm.</p><p>The test and the road are different places. Only one of them has a patient in the back.</p><h2>AI keeps passing, which tells us less, not more.</h2><p>The scoreboard looks one-sided. Back in 2024, AI was already clearing 80% on broad knowledge exams. By 2026 it was near 80% on harder tests too, including graduate-level science and fixing real bugs in real code.</p><p>The easy read is &#8220;the gap is closing.&#8221;</p><p>But there is an old rule here. Once a test becomes the goal, people aim at the test instead of the skill, and the test stops measuring anything real. Economists call it Goodhart&#8217;s Law.</p><p>A test everyone passes has stopped being a test. So the more AI aces them, the less the scores tell us about what people are for.</p><h2>A priest and a tech crowd reached the same verdict.</h2><p>In May 2026, Pope Leo XIV put it plainly:</p><blockquote><p>&#8220;They may imitate language, behavior, and analytical skills, or even simulate empathy and understanding, but they do not understand what they produce.&#8221;</p></blockquote><p>AI can copy how we talk. It can act caring. It still does not know what it is saying.</p><p>Here is the part I keep coming back to. The tech world reached the same place from the opposite side. Their version: AI does the easy part fast, so the hard part, the judgment, lands back on a person. The work does not vanish. It moves.</p><p>A priest and a tech podcast do not usually agree on anything. When they do, it is not a coincidence. It is a finding.</p><h2>Some of what you would automate is how your team learns.</h2><p>The Pope&#8217;s next point is the one most roadmaps miss:</p><blockquote><p>&#8220;humanity flourishes not despite limitations, but often through them.&#8221;</p></blockquote><p>In plain terms: the hard parts are often how we grow.</p><p>A job works the same way. Some struggle is just tedious, and a machine should take it. Some struggle is how a person becomes good at the work. Automate all of it, and the work ships faster while your people quietly get worse.</p><p>This argument has one real failure point, so I will name it. If AI gains lasting memory, a body, and real stakes in long relationships, the line blurs. Watch that line. But notice what would cross it: lived experience, not a higher score.</p><h2>Before you automate a job, ask what the test never measured.</h2><p>Stop scoring on &#8220;can AI beat the human.&#8221; Ask three smaller questions about each thing you plan to automate:</p><ol><li><p><strong>Where does a person have to understand the output, not just pass it on?</strong> That step is where your risk lives.</p></li><li><p><strong>What are you giving up: judgment, a hard-won skill, a human touch?</strong> Decide that on purpose, before a postmortem decides it for you.</p></li><li><p><strong>Which hard parts teach your team the most?</strong> Protect those. Automate the busywork.</p></li></ol><p>And one rule for the calls you cannot take back: if you cannot see how AI reached its answer, do not give it the final say.</p><p>AI aced the written test. It was always going to.</p><p>Whether to hand it the keys was never on the test.</p><p>That part is still yours.</p><h2>Source Notes</h2><ul><li><p><strong>The encyclical and its exact words.</strong> <em>Magnifica Humanitas</em>, Pope Leo XIV, issued 15 May 2026, &#8220;On Safeguarding the Human Person in the Time of Artificial Intelligence.&#8221; Paragraphs 99 and 118 are quoted word for word; the irreversible-decision rule is grounded in paragraph 198. External source: <a href="https://www.vatican.va/content/leo-xiv/en/encyclicals/documents/20260515-magnifica-humanitas.html">https://www.vatican.va/content/leo-xiv/en/encyclicals/documents/20260515-magnifica-humanitas.html</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Taste Tax]]></title><description><![CDATA[AI made execution cheap. Now judgment is the bottleneck.]]></description><link>https://www.producttheatre.com/p/the-taste-tax</link><guid isPermaLink="false">https://www.producttheatre.com/p/the-taste-tax</guid><dc:creator><![CDATA[Cam]]></dc:creator><pubDate>Thu, 04 Jun 2026 11:14:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zbAK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zbAK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zbAK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!zbAK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!zbAK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!zbAK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zbAK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:951408,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://producttheatre.substack.com/i/200433973?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zbAK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!zbAK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!zbAK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!zbAK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa7a8dd0b-8b9b-40c7-95b9-9a33b43f716c_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>AI did not remove the need for product judgment.</p><p>It removed the excuse for not having any.</p><p>When agents can produce five flows, three prototypes, two launch plans, and a passable landing page by Friday, the hard question changes. It is no longer &#8220;Can we build it?&#8221; It is &#8220;Which version is worth shipping?&#8221;</p><p>That scarce work is taste.</p><p>By taste, I do not mean aesthetics. I mean product discrimination: the ability to tell useful from noisy, sharp from generic, finished from merely working, and necessary from extra.</p><p>The taste tax is the bill that comes due when execution gets cheap.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.producttheatre.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.producttheatre.com/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>When building gets cheap, choosing gets exposed</h2><p>At Notion, the AI shift did not start with designers making prettier static mocks.</p><p>It moved the work into a shared prototype playground.</p><p>Brian Lovin, a designer on Notion AI, built an environment where designers could turn ideas into working prototypes with Claude Code. The point was not that every designer should become an engineer. The point was that AI products are hard to judge from static screens.</p><p>You have to feel the system behave.</p><p>A prototype that answers can still be wrong. It can use the wrong tone. It can hide the moment where trust breaks. It can complete the task and still feel generic in the user&#8217;s hands.</p><p>That is the right scene for the taste tax.</p><p>AI made the first version cheaper. It did not decide whether the interaction was good.</p><p>The new bottleneck is not making the artifact.</p><p>It is judging it while there is still time to change it.</p><h2>Taste is not decoration</h2><p>Taste is often treated as a soft word. That makes it easy to ignore.</p><p>Do not ignore it.</p><p>Taste is the operating standard underneath product work:</p><ul><li><p>What should we cut?</p></li><li><p>What should we polish?</p></li><li><p>What should we call this?</p></li><li><p>Which version helps the user think less?</p></li><li><p>Where does the feature cross from useful into trying too hard?</p></li><li><p>What would make this feel like us if the logo disappeared?</p></li></ul><p>Those are not decoration questions. They are product questions.</p><p>Before AI, weak taste was partially hidden by execution cost. Building was expensive, so fewer bad ideas made it all the way to users.</p><p>Now the gate is lower.</p><blockquote><p>Bad taste ships faster too.</p></blockquote><h2>The failure mode is generic abundance</h2><p>The first-order effect of AI is more output.</p><p>More drafts. More screens. More tests. More strategy docs. More prototypes. More &#8220;pretty good&#8221; versions of almost everything.</p><p>That sounds like leverage. Sometimes it is.</p><p>But abundance has a failure mode: teams stop discriminating. The roadmap only adds. The interface gets busier. The product starts to look like every other product built with the same models, the same templates, and the same unwillingness to say no.</p><p>This is the taste tax in the wild:</p><ul><li><p>The team can name twenty things to build and nothing to remove.</p></li><li><p>&#8220;The model wrote it&#8221; becomes a defense.</p></li><li><p>Code review catches bugs but not mediocrity.</p></li><li><p>The definition of done ends at &#8220;it works.&#8221;</p></li><li><p>Nobody can name a product they admire and explain the standard behind it.</p></li></ul><p>The product is not broken.</p><p>That is the problem. Broken work gets fixed. Generic work gets shipped.</p><h2>The signal is what the team refuses</h2><p>The easiest way to see taste is to ask what the team refused.</p><p>A team with taste can show you the versions they rejected. They can explain why a feature was removed, why a label changed, why one interaction got another hour and another got killed.</p><p>A team without taste can only show you the output.</p><p>This is why &#8220;move fast&#8221; became more dangerous in the agent era. Speed without discrimination converges to the mean. If everyone has access to similar generation tools, the difference is not who can produce more.</p><p>The difference is who can choose better.</p><p>James Bessen&#8217;s work on automation points to the broader pattern: technology often shifts work rather than simply removing it. When the task gets cheap, the judgment around the task becomes more valuable.</p><p>The artifact got cheaper.</p><p>The standard did not.</p><h2>Install a quality bar, not a taste lecture</h2><p>Do not tell teams to &#8220;have better taste.&#8221; That is not a process.</p><p>Make the standard visible at the moment work is about to ship.</p><p>Use one review with three questions:</p><ul><li><p>What did we refuse? Tests whether the team can cut. A good answer names a rejected version, feature, message, workflow, or interaction &#8212; with a clear reason.</p></li><li><p>What did we elevate? Tests whether the team improved beyond &#8220;it works.&#8221; A good answer names one change that made the output clearer, simpler, safer, more useful, or more distinctive.</p></li><li><p>What standard did we apply? Tests whether taste is shared or personal. A good answer is a reusable principle the next team can apply without the original reviewer in the room.</p></li></ul><p>This keeps the advice industry-agnostic. The artifact might be a product flow, pricing change, policy answer, data workflow, onboarding moment, service script, or internal tool. The standard is the same: the team can explain why this version deserves to exist.</p><p>Over time, save the best answers. That becomes the team&#8217;s taste library: not a mood board, but a record of good decisions.</p><h2>What changes Monday</h2><p>Pick one AI-assisted feature your team is about to ship.</p><p>Do one thing before launch: run a refusal review.</p><p>The launch decision should not be: &#8220;Does it work?&#8221;</p><p>It should be:</p><p><strong>Can the team explain what it refused, what it elevated, and what standard it used?</strong></p><p>If the answer is mostly silence, do not call the work finished.</p><p>Call it generated.</p><p>The market will not punish you because your team is slow.</p><p>It will punish you because everyone is fast now, and your product looks like theirs.</p><h2>Sources</h2><ul><li><p>Lenny&#8217;s Newsletter &#8212; I haven&#8217;t written a single line of code in six months: https://www.lennysnewsletter.com/p/i-havent-written-a-single-line-of</p></li><li><p>Lenny&#8217;s Newsletter &#8212; Why cultivating agency matters more than cultivating skills: https://www.lennysnewsletter.com/p/why-cultivating-agency-matters-more</p></li><li><p>James Bessen &#8212; Toil and Technology: https://www.imf.org/external/pubs/ft/fandd/2015/03/bessen.htm</p></li><li><p>Paul Graham &#8212; Taste for Makers: https://www.paulgraham.com/taste.html</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.producttheatre.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Product Theatre! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Straw Prompting]]></title><description><![CDATA[A demo is not an eval. That is the AI rollout mistake hiding in plain sight.]]></description><link>https://www.producttheatre.com/p/straw-prompting</link><guid isPermaLink="false">https://www.producttheatre.com/p/straw-prompting</guid><dc:creator><![CDATA[Cam]]></dc:creator><pubDate>Thu, 04 Jun 2026 11:03:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!jUTo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jUTo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jUTo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!jUTo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!jUTo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!jUTo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jUTo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png" width="1376" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1376,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1060924,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://producttheatre.substack.com/i/200433382?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jUTo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png 424w, https://substackcdn.com/image/fetch/$s_!jUTo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png 848w, https://substackcdn.com/image/fetch/$s_!jUTo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png 1272w, https://substackcdn.com/image/fetch/$s_!jUTo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0f6f3e97-d4d5-4f77-9a97-24995db7a785_1376x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg role="img" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><title></title><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p> </p><p>An AI demo is not an eval.</p><p>That is the mistake behind a lot of brittle enterprise AI rollouts. The demo proves the model can answer one polished question in one controlled room. An eval &#8212; a test set of real cases, scored against a written standard &#8212; tells you whether the system can survive actual users.</p><p>Most teams still ship the demo.</p><p>That is what I am calling <strong>straw prompting</strong>: building an AI feature out of a hand-tuned prompt, testing it on the cases that already work, and launching before the messy cases arrive.</p><p>It looks like a house.</p><p>It is not ready for weather.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.producttheatre.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.producttheatre.com/subscribe?"><span>Subscribe now</span></a></p><p></p><h2>The happy path is not the product</h2><p>Air Canada learned this in public.</p><p>In 2022, a customer used Air Canada&#8217;s website chatbot while booking travel after a death in the family. The chatbot told him he could buy a full-price ticket, travel, and then apply later for the airline&#8217;s bereavement fare.</p><p>That was wrong.</p><p>The airline&#8217;s actual bereavement policy said the discount could not be requested after travel had already happened. The chatbot even linked to the page with the correct policy, but the answer it gave in the conversation was still misleading.</p><p>The customer relied on the chatbot, travelled, and later asked for the fare difference. Air Canada refused. In 2024, British Columbia&#8217;s Civil Resolution Tribunal found Air Canada liable for negligent misrepresentation and ordered it to pay damages, interest, and fees.</p><p>The product lesson is not &#8220;chatbots are bad.&#8221;</p><p>It is sharper than that.</p><p>Air Canada had the correct information on its own website. The failure was that the interactive answer surface did not reliably behave like the policy it was supposed to represent.</p><p>That is the straw-prompting pattern in the wild.</p><p>The company did not just need a chatbot that could answer bereavement questions. It needed a system that could notice policy conflict, ground answers in the current source of truth, and escalate when the answer affected money, rights, or obligations.</p><p>A working answer was not enough.</p><p>The answer had to be governed like part of the product.</p><h2>The wolf is variance</h2><p>The wolf is not a smarter model from a competitor.</p><p>The wolf is variance: the messy spread of real-world inputs that your demo never saw.</p><p>Users misspell things. They paste half a policy. They use internal acronyms. They ask two questions at once. They omit context. They ask in anger. They ask in formats your prompt writer did not imagine.</p><p>OpenAI&#8217;s eval guidance makes the point plainly: AI systems need task-specific tests that reflect real-world conditions, edge cases, and continuous change. &#8220;Looks good to me&#8221; is an anti-pattern, not a launch gate.</p><blockquote><p>That is the trap. A demo that works ten times in a row does not tell you how the system behaves on the eleventh thousand.</p></blockquote><h2>The launch review should expose the failure mode</h2><p>You do not need a 40-point governance checklist to spot straw prompting.</p><p>You need a launch review that separates symptoms from causes. Run it against five patterns:</p><ul><li><p>Polished examples, not error rates &#8212; the team tested the demo, not the system. Require a real-input eval set before launch.</p></li><li><p>No one can name the top failure modes &#8212; the risk model is missing. Write the failure list before tuning the prompt.</p></li><li><p>Output quality judged by vibes &#8212; the team has no shared standard. Create a rubric before reviewing answers.</p></li><li><p>&#8220;Guardrails later&#8221; in the plan &#8212; controls are being treated as cleanup. Define refusal, escalation, and source-of-truth rules now.</p></li><li><p>The prompt writer also wrote the tests &#8212; the eval is likely overfit to the happy path. Add messy cases from real usage, edge conditions, policy boundaries, and expert review.</p></li></ul><p>The point is not to punish the team for moving quickly. It is to find the root cause while the cost of fixing it is still low.</p><h2>Build the test before the prompt</h2><p>The fix is not a giant governance program. It is a different order of operations.</p><p>Start with the test.</p><p>For any AI feature, collect 100 to 300 real or realistic inputs from the hardest part of the workflow. Include short requests, long requests, ambiguous requests, boundary cases, policy conflicts, and cases where the right answer is &#8220;I do not know.&#8221;</p><p>Then write the rubric.</p><p>What counts as correct? What counts as unsafe? When should the model refuse? When should it escalate to a human? What evidence should it cite? What must it never invent?</p><p>Only then tune the prompt.</p><p>The prompt has to earn its way through the mess. Otherwise you are not improving the product. You are decorating the demo.</p><h2>The brick house has gates, owners, and reruns</h2><p>Once the root cause is clear, prevention is boring on purpose.</p><p>Brick-house AI rollouts put four controls in place before launch:</p><ol><li><p><strong>A golden set.</strong> The small, trusted collection of examples that defines what &#8220;good&#8221; looks like. It includes normal cases, edge cases, and failures you never want to see twice.</p></li><li><p><strong>A named owner.</strong> Someone owns the eval set, the rubric, and the launch threshold. Without an owner, quality becomes a meeting mood.</p></li><li><p><strong>Escalation rules.</strong> The system knows when to refuse, when to cite the source of truth, and when to send the user to a person.</p></li><li><p><strong>Scheduled reruns.</strong> Models change. Prompts change. Retrieval changes. Product policy changes. The system that passed Tuesday may not pass next Tuesday.</p></li></ol><h2>What changes Monday</h2><p>Pick the next AI feature waiting for launch approval.</p><p>Do one thing: replace the demo review with an eval review.</p><p>The launch decision should not be: &#8220;Did the examples look good?&#8221;</p><p>It should be:</p><p><strong>Did the system pass the real cases, against the written rubric, with clear escalation rules and an owner for reruns?</strong></p><p>If not, do not approve the launch.</p><p>Approve the demo, if you want.</p><p>Just do not mistake it for a brick house.</p><h2>Sources</h2><ul><li><p>OpenAI &#8212; Evals: https://platform.openai.com/docs/guides/evals</p></li><li><p>OpenAI Cookbook &#8212; Evals design guide: https://cookbook.openai.com/examples/evaluation/evals_design_guide</p></li><li><p>NIST &#8212; AI Risk Management Framework: https://www.nist.gov/itl/ai-risk-management-framework</p></li><li><p>Deeth Williams Wall &#8212; BC Tribunal Finds Air Canada Liable: https://www.dww.com/articles/bc-tribunal-finds-air-canada-liable-for-inaccurate-advice-given-by-website-chatbot</p></li><li><p>American Bar Association &#8212; BC Tribunal Confirms Companies Remain Liable: https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/</p><p></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.producttheatre.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Product Theatre! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>