Przejdź do treści

How I AI

Claire Vo
How I AI
Najnowszy odcinek

113 odcinków

  • How I AI

    Jev for beginners: how to use it and what to build

    28.09.2026 | 26 min.
    Jev is TypeSafe AI’s new decision model. It returns type-safe structured values (a choice, a score, a probability) instead of generated text, at 4 cents per million input tokens with no output charge. This week I ran it on five real projects: PR categorization, a meta-analysis of my own Claude and Codex sessions, Gmail triage, the ChatPRD product insights graph, and a live audience dashboard built from 4,500 YouTube comments.

    What you’ll learn:
    What makes Jev fundamentally different from every other model I’ve used
    How I analyzed 1,700 PRs for 9 cents and what I found out about where my engineering effort actually went
    The personal meta-analysis you can run on your own Claude and Codex sessions right now
    Why I stopped using Jev alone, and what I pair it with now
    How I turned 4,500 YouTube comments into a searchable audience dashboard for almost nothing
    The real-time app I built in an afternoon that shows something surprising about Jev’s speed
    Why Jev’s pricing model is different from any LLM I’ve used, and what it makes practical to build
    The ChatPRD product insights project: 1,100 signals, 200,000 classifications, and what it cost me
    —
    Brought to you by:
    OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more
    —
    In this episode, we cover:
    (00:00) Jev launch and what makes it different from every other model
    (02:49) Type-safe values explained
    (05:28) Understanding Jev outputs
    (07:39) Use case 1: PR categorization and pairwise clustering
    (11:12) Use case 2: analyzing your own local Claude Code and Codex sessions
    (13:00) Use case 3: Gmail triage with Jev scoring and LLM follow-up
    (14:30) Use case 4: ChatPRD’s product insights graph
    (18:17) Demo: How I AI audience signal dashboard
    (22:14) Demo: voice-to-color emotion-mapping app
    (25:16) Jev week recap and what’s coming in episode 2
    —
    Tools referenced:
    • Jev (TypeSafe AI): https://typesafe.ai
    • Vercel: https://vercel.com/ai
    • GitHub API: https://docs.github.com/en/rest
    • YouTube Data API v3: https://developers.google.com/youtube/v3
    • OpenAI Realtime Voice API: https://platform.openai.com/docs/guides/realtime
    • Gemini 3.5 Flash-Lite: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite
    • API Ninjas Quotes API: https://api-ninjas.com/api/quotes
    —
    Where to find Claire Vo:
    ChatPRD: https://www.chatprd.ai/
    Website: https://clairevo.com/
    LinkedIn: https://www.linkedin.com/in/clairevo/
    X: https://x.com/clairevo
    —
    Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.
  • How I AI

    Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?

    22.09.2026 | 38 min.
    I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still has me split. There’s a creative result I got completely wrong, an LLM judge that disagreed with me, and a return to Barbie Bench: the 3D fashion game that keeps reminding me how far we have to go. The hands are tragic. AGI has not arrived.

    What you’ll learn:
    How I run the How I AI bench blind, and what gets an output a bad score before I even know which model made it
    Why Astra won my heart while Opus 5.5 might be overall strongest, especially for long-running agents and B2B frontend
    Where Sol still wins me over on clear writing, readable PRDs, and price
    The character SVG results that completely overturned my prediction about Anthropic
    What happened when I asked these models to edit video, and why I think skills explain part of the disappointment
    Why an LLM judge disagreed with my rankings, and what it was rewarding that I wasn’t
    —
    In this episode, we cover:
    (00:00) LIVE setup and new model launches
    (01:30) What’s new in Opus 5.5, Sol, and Luna
    (04:11) Guardrails, personality, and speed
    (09:00) The How I AI bench and blind evaluation process
    (11:31) Email and personal-productivity results
    (13:50) Frontend prototype vibe checks
    (24:10) Backend, agent personality, and long-running tasks
    (28:25) SVG illustration test
    (29:48) AI video-editing results
    (30:43) Predictions before the reveal
    (31:20) Barbie Bench: the 3D fashion-game test
    (34:17) Results: Astra, Sol, and Opus 5.5
    (35:04) Writing clarity and creative surprises
    (36:51) Why the LLM judge disagreed with me
    (37:24) What each model is actually best for
    —
    Tools referenced:
    • Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5
    • GPT-6 Sol and Luna: https://openai.com/index/introducing-gpt-6-sol-and-luna/
    • Codex (OpenAI): https://openai.com/codex
    —
    Where to find Claire Vo:
    ChatPRD: https://www.chatprd.ai/
    Website: https://clairevo.com/
    LinkedIn: https://www.linkedin.com/in/clairevo/
    X: https://x.com/clairevo
    —
    Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.
  • How I AI

    I left Claude for months. Opus 5.5 is why I'm back

    22.09.2026 | 24 min.
    I’ve been off Claude for months. Not because it got dumb, but because it got annoying. The rambling, the hedging, the preachy little disclaimers on tasks that didn’t need them. I moved most of my daily work to Codex and I didn’t miss it. Then Anthropic shipped Opus 5.5: 40% cheaper than Opus 5, faster, and with what they’re calling a fundamentally different alignment approach. I ran it for a week across real work, including four long-running agentic tasks, a full ChatPRD homepage redesign, an SVG benchmark, and one very firm refusal, and I’m ready to give you the honest verdict. There’s a lot to like. There are still two things that drive me a little crazy. And there’s one capability I genuinely wasn’t expecting.

    What you’ll learn:
    Why I walked away from Claude entirely, and what it took for me to come back
    The real cost math on Opus 5.5 and why pricing matters more for agentic work than single prompts
    What happened when I ran four long-running agentic tasks, including one that tried to manipulate Claude mid-run
    Why Opus 5.5 is now my go-to for frontend prototyping, and where it still lets me down
    The one capability I genuinely didn’t see coming, and no other model in my stack can match it
    The moment Opus 5.5 told me flat-out no, and what that says about where Anthropic’s safety posture actually lands in practice
    Where Codex still wins, and how I’m splitting my model stack after a full week of testing
    —
    In this episode:
    (00:00) Why I stopped using Claude
    (01:02) What Anthropic says Opus 5.5 is
    (01:54) Cost, speed, and benchmark overview
    (03:20) Safety, alignment, and the cybersecurity limits
    (05:02) How I AI bench
    (05:39) Voice test: is it actually not annoying?
    (07:54) Long-running agentic task results
    (10:50) Frontend prototyping
    (17:23) Writing voice and email
    (19:41) SVG illustrations
    (20:46) Video editing
    (21:42) My verdict: what it’s good at, what it still isn’t
    —
    Tools referenced:
    • Claude Opus 5.5: https://www.anthropic.com/claude-opus-5-5
    • ElevenLabs MCP connector: https://elevenlabs.io/mcp
    • Codex (OpenAI): https://openai.com/codex
    —
    Where to find Claire Vo:
    ChatPRD: https://www.chatprd.ai/
    Website: https://clairevo.com/
    LinkedIn: https://www.linkedin.com/in/clairevo/
    X: https://x.com/clairevo
    —
    Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.
  • How I AI

    How Warp ships 2,000 PRs a month with AI factories | Zach Lloyd (CEO, Warp)

    21.09.2026 | 46 min.
    Zach Lloyd is the co-founder and CEO of Warp, an AI-powered terminal and software factory platform used by tens of thousands of engineers. Before Warp, he spent nearly a decade at Google, including time as a principal engineer on Google Sheets. He built Warp from the ground up as a modern, AI-native alternative to legacy terminals, and the team has since expanded into software factories: a full cloud-based system that takes an idea in Slack all the way through to a merged PR.

    In this episode:
    Why a software factory is more than a coding agent
    The public Slack → Linear → GitHub → QA workflow
    Human interactions per PR as a signal of automation and throughput
    Why human review is still the bottleneck
    Scoring agent runs, finding failure modes, and self-improving agent workflows
    Replaying real tasks to choose model cost and quality tradeoffs
    CEO workflows with Figma MCP, Granola, and research agents
    —
    Brought to you by:
    DX—Engineering intelligence for the AI era
    OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more
    —
    In this episode, we cover:
    (00:00) Intro
    (02:35) Warp’s AI software factory, Wilson
    (09:23) Automatic factory triggers
    (11:12) The engineering leader dashboard Zach wishes he’d had
    (15:18) How code review is changing in an AI factory
    (17:08) Tracking cost per PR across model configs
    (18:47) Using LLM-as-a-judge to score every agent run
    (20:02) Catching redundant tests
    (22:19) How the factory self-improves from failed runs
    (26:03) Quick recap
    (28:33) Building a cost-quality Pareto chart for model selection
    (31:35) How Zach uses AI for non-technical CEO work
    (32:10) Figma MCP demo
    (35:43) Granola MCP demo
    (36:41) GOG CLI demo
    (38:20) Thinking in parallel tasks instead of sequential ones
    (40:42) Zach’s prompting strategy for factory tasks
    (44:48) Where to find Zach
    —
    Tools referenced:
    • Warp (AI terminal and software factories): https://warp.dev
    • Warp Factories: https://warp.dev/factories
    • Linear (project and issue tracking): https://linear.app
    • GitHub (version control and PR management): https://github.com
    • Slack (team communication and factory input layer): https://slack.com
    • Sentry (crash reporting and automated issue triggers): https://sentry.io
    • Figma (design, used via Figma MCP): https://figma.com
    • Granola (AI meeting notes and MCP integration): https://granola.so
    • Grok Bot (fast inference, cost/quality trade-off): https://x.ai/bot/guides/grok-bot-101
    —
    Where to find Zach:
    X: https://x.com/ZachLloydTweets
    —
    Where to find Claire:
    ChatPRD: https://www.chatprd.ai/
    Website: https://clairevo.com/
    LinkedIn: https://www.linkedin.com/in/clairevo/
    X: https://x.com/clairevo
    —
    Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.
  • How I AI

    Muse review: The personal AI agent that gets consumer UX right

    16.09.2026 | 37 min.
    I spent a few hours putting Meta’s Muse, its new personal AI agent, through a real first-pass test: onboarding, calendar management, goal setting, a one-shot family morning newsletter, browser-based shopping, and the animated avatar that honestly surprised me.

    What you’ll learn:
    Why Muse is the best-designed personal agent I’ve tested, and what specifically made it feel that way
    The one-shot family PDF Muse produced that Claude and Codex never quite nailed
    How Muse’s permission model works, and why it’s different from every other agent I’ve used
    Why I set up a sleep training goal in Muse, and what it revealed about agent tone
    The activity feed feature I immediately wished Codex and Claude Code had
    Where Muse failed, and what it says about the limits of this category right now
    The animated avatar decision that showed me what top-of-craft AI product design actually looks like
    —
    Brought to you by:
    Optimizely—Your AI agent orchestration platform for marketing and digital teams
    OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more
    —
    In this episode, we cover:
    (00:00) What Muse is and who it’s actually built for
    (04:41) Signing in and the onboarding flow
    (07:16) The activity feed and its task lineage
    (08:24) First real task: managing the family calendar and deleting soccer practice
    (09:48) Requesting a morning newsletter PDF
    (14:29) The personalized news feed and how I set it up
    (16:05) The “Ideas” feature as an out-of-the-box prompt library
    (17:10) Setting up personal goals (water, shoes, and sleep training)
    (21:40) Library: documents, websites, images, videos, and podcasts
    (23:11) Quick recap and what I love
    (23:56) Activity feed design deep dive: tool calls and step-by-step lineage
    (25:18) How Muse handles permissions
    (26:12) The animated avatar: Polly becomes Slime, the teal dragon
    (29:34) Browser use test: shopping for New Balance 9060s (not great)
    (31:15) Browser use test 2: buying IMAX tickets for The Odyssey (much better)
    (33:34) TL;DR and what I’ll actually use Muse for going forward
    —
    Tools referenced:
    • Muse: https://muse.ai/
    • Stripe Link (payment method featured in Muse): https://link.com
    • 1Password (future Muse integration mentioned): https://1password.com
    • OpenClaw (Claire’s previous personal agent setup): https://openclaw.ai/
    • Grok Bot (Grok-based agent from prior stack): https://x.ai/news/introducing-grok-bot
    • Codex (OpenAI coding agent, comparison point): https://openai.com/codex
    • NotebookLM (Google, comparison to Muse’s podcast generation): https://notebooklm.google.com
    —
    Where to find Claire Vo:
    ChatPRD: https://www.chatprd.ai/
    Website: https://clairevo.com/
    LinkedIn: https://www.linkedin.com/in/clairevo/
    X: https://x.com/clairevo
    —
    Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.
Więcej Technologia podcastów
O How I AI
How I AI, hosted by Claire Vo, is for anyone wondering how to actually use these magical new tools to improve the quality and efficiency of their work. In each episode, guests will share a specific, practical, and impactful way they’ve learned to use AI in their work or life. Expect 30-minute episodes, live screen sharing, and tips/tricks/workflows you can copy immediately. If you want to demystify AI and learn the skills you need to thrive in this new world, this podcast is for you.
Strona internetowa podcastu

Słuchaj How I AI, Technologicznie i wielu innych podcastów z całego świata dzięki aplikacji radio.pl

Uzyskaj bezpłatną aplikację radio.pl

  • Stacje i podcasty do zakładek
  • Strumieniuj przez Wi-Fi lub Bluetooth
  • Obsługuje Carplay & Android Auto
  • Jeszcze więcej funkcjonalności