ANALOG / AIE 2026 / Slides worth keeping The Signal · Slides worth keeping
The slides worth keeping Four days at the AI Engineer World's Fair, cut to the slides worth keeping, the arguments, the frameworks, the numbers, and the one-liners that defined the week, plus the rooms they filled, shot from the seats.
49 photographs · Jun 29 - Jul 2, 2026 · Photographs from the seats by Matt Montenegro
Day 1 · Jun 29 The SDLC eating its own tail: production makes signal, signal becomes fixes, and fixes become the next production. Promotion as memory: a proven workflow is distilled into a reusable SKILL.md, and the raw recipe is retired so the agent stops paying attention twice. Quality as Swiss cheese: automated evals, human transcript review, and production monitoring, each layer catching what the last one missed. One memory for many agents, updated in real time, then ‘dreamed’ into shape between sessions. The takeaway of the week on one slide: the harness is the product, and traces → evals → the loop keep improving it. The improvement loop in six boxes, run, trace, eval, cluster, fix, until the failed trace becomes a regression test. Three ways to build, drawn as toy cars: waterfall waits for the whole car, agile ships a skateboard then a bike, and the AI row keeps the wheels turning. Small enough to learn before you build the whole car. The gap benchmarks miss: the agent's code passes the test but ignores the repo's own formatMoney(), and the number you optimize isn't the number your users feel. Day 2 · Jun 30 Three agents in one product: an engagement agent on the floor, a workflow agent, and a development agent. The shift underneath the year: stop decomposing tasks, start handing the model the context to decide like you would. One overnight prompt hands off the whole feature, tests, PRs, review, and asks only for a report in the morning. The honest asterisk on agent speed: “My code has annoyed my engineers. It has caused SEVs.” Said plainly from the main stage: building is part of the job now. The forward-deployed job, layer by layer: DevOps, then data integration, then custom solutions, then enablement. Forward-deployed engineering is dead; long live forward-deployed engineering, the week’s most-repeated paradox. A case study does the arguing: one system for chat and voice, 70% resolution and support costs down 60%. The whole forward-deployed playbook on two lines: always be scoping, and scale with tokens. The job description for what comes next: harnesses, evals, context, skills, and taste on agent output. Six people, not thirty: Amazon’s Bedrock Mantle, built with Kiro, as the pathfinder for agent-native teams. One big takeaway from an Amazon engineer: intentionally change the way you work. Pillars of AI governance, read against a wall of the fair's sponsors. The security slide everyone screenshotted: private data plus untrusted content plus external comms is the lethal trifecta. The wall every engineering leader hits, now moved: fragmented time became buildable, and a coding block at each end of the day is enough to ship real things. The forward-deployed job description as a joke that isn't: eight years as a staff engineer, six years of client-facing sales, four as a solution architect, and deep fluency in half a dozen stacks. Every manual step becomes one the next customer skips: 80% less custom engineering per agent once the delivery bottleneck turns into a self-serve roadmap. The asterisk on 'built with only six people': two distinguished engineers, one senior principal, and three principals, shipping a whole inference data plane on Kiro. The five habits of frontier teams on one slide: invest in agent context, slow down to speed up, feed agents don't babysit, make intent explicit, and shift testing left. What's still hard, named honestly: burnout, FOMAT (the fear of missing agent time), and reviewing AI output that is sometimes harder than writing it yourself. Always be scoping, rendered literally: a workhorse in a green field with rocket boosters strapped to all four legs. Day 3 · Jul 1 Agents need enterprise context and constraints to be ‘verification aware’, architecture, standards, and guardrails, fed in or trained in. The harness becomes a discipline: keep an implementation-notes.md, log the deviations, keep going. A history lesson before the hype: cybernetics in the 1940s, and yes, it shares a Greek root with Kubernetes. The long road to the LLM: from 1998 to today, gated for years by ‘not enough compute’ and ‘not enough data’. 1972: the first videogame tournament, played on a research machine at Stanford’s AI Lab. A riff on Andreessen: software ate the world, but verifiable simulation is what eats the physical one. One skill, every style: a design skill that bends to any brand instead of flattening it. A design skill with a manifesto: one that refuses to look AI-generated. Evals at DeepMind scale: machine-readable checks, proof of skill-lift before a merge. Where the field got its name: the 1956 Dartmouth Summer Research Project, the workshop that first called it artificial intelligence. Learning a policy directly in the world: a row of robot arms collecting grasping episodes around the clock with autonomous self-supervision. One talk's north star, drawn as a stick figure at a telescope: embodied artificial general intelligence. One skill, twenty-plus themes: the same brief rendered as Cold Snap, Off-Register, Press Quaternary, and a dozen more, no two alike. The takeaway for taste: save your design preferences in an AGENTS.md so the agent stops guessing what you like. Day 4 · Jul 2 Making agent work visible: the same artifact, in Slack and in the app. The cautionary tale of the week, in a headline: an AI agent wiped a company’s entire database in nine seconds. The main hall from the seats: the fair's sponsor wall glowing on the left, a talk on why AI workloads live or die on network latency, and a full house watching. The audience poll that framed the week: evaluation narrowly tops the list of hardest problems, with orchestration, inference, and security a few points behind and only 4% seeing no challenge at all. Lesson zero, read over two attendees' shoulders: be model and harness agnostic, because the best one changes weekly and the token vendor's incentives are not yours. The number the room came for: 99% of PRs agent-generated and 100% human plus agent reviewed, tallied across Claude Code, Codex, Cursor, and a wall of other tools.