Multimodal AI: When Models See, Hear, and Speak
Aaron Baughman on the rise of multimodal AI — vision models, audio processing, and the path to autonomous digital workers that interact with the world.
Aaron Baughman
AI Engineer & Master Inventor
IBM
Dr. Sarah Chen
Host & AI Research Lead
Former DeepMind researcher with a PhD in Machine Learning from Stanford. Covers AI, quantum, and computational breakthroughs.
About This Episode
In Episode 128 of The Frontier Tech Show, host Dr. Sarah Chen sits down with Aaron Baughman, AI Engineer & Master Inventor at IBM, to discuss "Multimodal AI: When Models See, Hear, and Speak." This ai podcast episode, published on March 18, 2026 as part of Season 4, runs 57:12 and covers scaling laws and compute trajectory, reasoning models and chain-of-thought, multimodal capabilities, and agent architectures, open vs closed source, training efficiency, data quality and curation, safety and alignment. The conversation provides a deep dive into the current state of ai technology, exploring both the technical breakthroughs driving the field forward and the real-world challenges that remain.
Aaron Baughman brings deep expertise to this conversation. As AI Engineer & Master Inventor at IBM, Aaron Baughman offers a front-line perspective on scaling laws and compute trajectory that goes beyond surface-level analysis. The discussion covers how ai has evolved over the past year, what the key inflection points have been, and where the technology is heading in the next twelve to eighteen months. Whether you are a practitioner, investor, or simply following the ai space, this episode delivers insights you will not find elsewhere.
Listeners will come away from this episode with a clear understanding of scaling laws and compute trajectory and its implications for the broader ai landscape. The conversation covers the science, the engineering, the economics, and the policy dimensions of multimodal ai: when models see, hear, and speak, making it essential listening for anyone who wants to understand where ai is going in 2026 and beyond.
Key Topics Discussed
- Scaling laws and compute trajectory: The discussion explores scaling laws and compute trajectory in depth, examining current capabilities, limitations, and the trajectory of development. Aaron Baughman shares specific examples and data points from work at IBM, giving listeners a concrete sense of where the technology stands today and what milestones to watch for.
- Reasoning models and chain-of-thought: The discussion explores reasoning models and chain-of-thought in depth, examining current capabilities, limitations, and the trajectory of development. Aaron Baughman shares specific examples and data points from work at IBM, giving listeners a concrete sense of where the technology stands today and what milestones to watch for.
- Multimodal capabilities: The discussion explores multimodal capabilities in depth, examining current capabilities, limitations, and the trajectory of development. Aaron Baughman shares specific examples and data points from work at IBM, giving listeners a concrete sense of where the technology stands today and what milestones to watch for.
- Agent architectures: The discussion explores agent architectures in depth, examining current capabilities, limitations, and the trajectory of development. Aaron Baughman shares specific examples and data points from work at IBM, giving listeners a concrete sense of where the technology stands today and what milestones to watch for.
- Open vs closed source: The discussion explores open vs closed source in depth, examining current capabilities, limitations, and the trajectory of development. Aaron Baughman shares specific examples and data points from work at IBM, giving listeners a concrete sense of where the technology stands today and what milestones to watch for.
- Training efficiency: The discussion explores training efficiency in depth, examining current capabilities, limitations, and the trajectory of development. Aaron Baughman shares specific examples and data points from work at IBM, giving listeners a concrete sense of where the technology stands today and what milestones to watch for.
- Data quality and curation: The discussion explores data quality and curation in depth, examining current capabilities, limitations, and the trajectory of development. Aaron Baughman shares specific examples and data points from work at IBM, giving listeners a concrete sense of where the technology stands today and what milestones to watch for.
- Safety and alignment: The discussion explores safety and alignment in depth, examining current capabilities, limitations, and the trajectory of development. Aaron Baughman shares specific examples and data points from work at IBM, giving listeners a concrete sense of where the technology stands today and what milestones to watch for.
Episode Details
Aaron Baughman on the rise of multimodal AI — vision models, audio processing, and the path to autonomous digital workers that interact with the world.
Episode Transcript
Full transcript of "Multimodal AI: When Models See, Hear, and Speak" — Episode 128 of The Frontier Tech Show with Aaron Baughman, AI Engineer & Master Inventor at IBM. (938 words)
COLD OPEN
Marcus Webb: Aaron, I've been following multimodal for a while, and I have to say — what's happened in the last year feels different. Not just incremental progress, but a qualitative shift. Am I reading that right?
Aaron Baughman: You are. And I think the reason it feels different is that we've crossed the threshold from 'interesting science' to 'practical technology.' That's a transition that many fields never make. The fact that we're talking about when models see, hear, and speak in terms of deployment timelines and unit economics, not just research papers — that's the signal.
Dr. Sarah Chen: Welcome to TechNova. I'm Dr. Sarah Chen.
Marcus Webb: And I'm Marcus Webb. Today we're joined by Aaron Baughman, AI Engineer & Master Inventor at IBM. Aaron, welcome to the show.
Aaron Baughman: Thanks for having me. Looking forward to this.
SEGMENT 1: The State of the Field
Dr. Sarah Chen: Aaron, for listeners who are new to this topic, can you explain what multimodal actually involves and why it matters?
Aaron Baughman: At its core, multimodal is about scaling laws and compute trajectory. That sounds simple, but the implications are profound. When you can do reasoning models and chain-of-thought reliably and at scale, it changes what's possible in AI. The applications range from multimodal capabilities to agent architectures, and we're just scratching the surface.
Marcus Webb: How did we get here? What was the path from idea to reality?
Aaron Baughman: It was a long path — decades, in some cases. The foundational research in open vs closed source goes back years, but it was always limited by training efficiency. What changed is that we solved that limitation — through a combination of better technology, better understanding, and honestly, better computing power. Once the bottleneck cleared, everything downstream accelerated.
Dr. Sarah Chen: And where are we now on that path?
Aaron Baughman: We're in the early deployment phase. The technology works. We're proving it in real-world conditions. The next challenge is scaling — making it cheaper, more reliable, and more accessible. That's an engineering challenge, not a science challenge, and engineering challenges are solvable with enough time and resources.
SEGMENT 2: The Technical Details
Marcus Webb: Aaron, I want to get into the technical details. What makes your approach different from what's been tried before?
Aaron Baughman: The traditional approach to multimodal relied on data quality and curation. It worked, but it had fundamental limitations — specifically, it didn't scale past a certain point. Our approach is different because we use safety and alignment to bypass those limitations entirely. Instead of trying to optimize within the old framework, we created a new framework.
Dr. Sarah Chen: What was the key insight that enabled that?
Aaron Baughman: It was actually a cross-disciplinary insight. Someone on our team had experience in vision, and they noticed a parallel between a problem in that field and our problem in multimodal. They brought a technique over, adapted it, and it worked. The biggest breakthroughs often come from the intersection of fields, not from deep within one field.
Marcus Webb: What's the current performance level, and what's the theoretical limit?
Aaron Baughman: We're currently at about 60 percent of what we believe is the theoretical limit. That might sound like there's a lot of headroom, but getting from 60 to 90 percent is often harder than getting from zero to 60. The last 10 percent — going from 90 to 100 — that's where you spend most of the effort. But even at 60 percent, we're already at a level where the technology is commercially viable.
SEGMENT 3: Real-World Impact
Dr. Sarah Chen: Let's talk about impact. Who benefits from this, and how?
Aaron Baughman: The impact is broad. In the near term, regulatory landscape is the primary application — and that alone justifies the investment. But the second-order effects are where it gets really interesting. Once you have multimodal working at scale, it enables things that weren't possible before — investment and compute infrastructure, new business models, new capabilities. It's a platform technology, not just a point solution.
Marcus Webb: What about the risks? What could go wrong?
Aaron Baughman: I take risks seriously, and there are real ones. scaling laws and compute trajectory at scale is untested — we're confident, but there could be surprises. There's the regulatory risk — if policymakers move too slowly, deployment stalls. And there's the societal risk — any transformative technology has distributional effects, and we need to be thoughtful about who benefits and who's displaced.
Dr. Sarah Chen: How do you think about the ethical dimensions?
Aaron Baughman: It's something we discuss internally a lot. The technology itself is neutral — it's a tool. But how it's deployed, who has access to it, what safeguards are in place — those are choices, and they matter. I think the tech industry as a whole needs to do a better job of engaging with these questions proactively, not reactively.
SEGMENT 4: Looking Forward
Marcus Webb: Aaron, what's your vision for where this field is in five years?
Aaron Baughman: In five years, I think multimodal will be unremarkable — and that's the goal. When a technology becomes unremarkable, it means it's become infrastructure. It's just part of how things work. That's what happened with the internet, with smartphones, with cloud computing. I think multimodal is on that same trajectory, and the five-year mark is when it crosses from 'exciting new technology' to 'standard tool that everyone uses.'
Dr. Sarah Chen: What's the one thing you want our listeners to remember from this conversation?
Aaron Baughman: That the future is being built right now, by people who are solving hard problems in labs and offices and factories. It's not science fiction — it's engineering. And engineering, when done well, is the most powerful force for progress that humanity has ever developed.
Marcus Webb: Aaron Baughman, AI Engineer & Master Inventor at IBM. Thank you for a really thought-provoking conversation.
Aaron Baughman: Thank you both. I loved this.
Dr. Sarah Chen: And thanks to all of you for listening. This is TechNova — see you next time.
Why This Episode Matters
This episode matters because ai is at a critical juncture in 2026. The conversation between Dr. Sarah Chen and Aaron Baughman cuts through the hype to deliver a grounded, evidence-based assessment of where scaling laws and compute trajectory actually stands. For decision-makers in technology, finance, and policy, understanding the nuances discussed here is essential for making informed bets on the future of ai.
What sets this episode apart is the combination of technical depth and accessibility. Aaron Baughman explains complex concepts in ai without oversimplifying, making this episode valuable for both experts and newcomers to the field. The discussion of scaling laws and compute trajectory and reasoning models and chain-of-thought alone makes this episode worth listening to, but the broader conversation about the future direction of ai technology is what makes it truly essential.