scripod.com

AI Agents Hit The Verification Wall

Shownote

The episode focused on practical AI workflow design, especially how Fable fits as a high-cost planning and audit model rather than a default execution model. The hosts discussed compound engineering, verification loops, Caveman-style terse prompting, and how AI work changes communication habits. They also covered Microsoft Frontier Co and the broader move toward embedded AI engineering for enterprises. The final news segment debated Wired’s report on Meta’s Project Cannes and whether aggressive safety testing belongs inside companies, with contractors, or under stronger oversight. Key Points Discussed 00:00:18 Episode Intro And Hosts 00:01:36 Weekend Fable Use Cases 00:05:56 Fable Audits For AI Workflows 00:09:20 Compound Engineering And Verification Loops 00:15:39 Using Fable As The Expert Model 00:19:32 Microsoft Frontier Co And Embedded Engineers 00:25:47 AI Audits And Working Worldviews 00:34:04 Caveman Plugin And Token Efficiency 00:38:14 Field Guide To Fable Unknowns 00:39:49 GPT-5.6, Watermelon And Codex Ultra 00:41:37 Claude Suggested Tasks And Branches 00:44:16 Meta Project Cannes Safety Testing 00:58:07 Fable Usage Credits Clarified The Daily AI Show Co Hosts: Karl Yeh, Beth Lyons, Brian Maucere, Andy Halliday

Highlights

This episode explores the strategic use of high-cost AI models like Fable for planning and auditing, rather than routine execution. The hosts discuss compound engineering, verification loops, and the shift toward embedded AI engineering in enterprises, alongside a debate on Meta's aggressive safety testing practices.
00:00
Fable analyzes complex legal and financial situations.
01:41
Confusion over Fable's usage credits and discount deadlines
08:12
Distilling high-end model reasoning into work plans
09:23
Building is cheap, verification is expensive.
15:40
Fable consumed many credits with poor results
25:09
Changing the worldview is the hardest part.
28:43
Expertise reduces the threat from AI
34:08
Caveman cuts token spend by 65%
38:14
Navigating unknown territory in AI
39:53
Ultra will be available in Codex
41:42
Helpful for delegating orthogonal tasks
44:18
Extreme stress testing of frontier models is necessary for safety.
58:08
Fable usage may end by July 7th

Chapters

Episode Intro And Hosts
00:00
Weekend Fable Use Cases
01:36
Fable Audits For AI Workflows
05:56
Compound Engineering And Verification Loops
09:20
Using Fable As The Expert Model
15:39
Microsoft Frontier Co And Embedded Engineers
19:32
AI Audits And Working Worldviews
25:47
Caveman Plugin And Token Efficiency
34:04
Field Guide To Fable Unknowns
38:14
GPT-5.6, Watermelon And Codex Ultra
39:49
Claude Suggested Tasks And Branches
41:37
Meta Project Cannes Safety Testing
44:16
Fable Usage Credits Clarified
58:07

Transcript

Karl Yeh: Hey, what's going on, everybody? Welcome to the Daily AI Show. Today is July 6, 2026, episode 761. Quickly making our way towards the back half of the 700s and the 800. But before we get to 800, we're going to be hitting our three-year anniversar...