Ivy: Structure your API calls just right, and 85% of your token costs could just... vanish.
Marcus: I'm Marcus.
Ivy: And I'm Ivy.
Marcus: And this is Broke It, your daily download on the bleeding edge of AI. It’s Wednesday, September 23rd.
Ivy: We've got a packed show for you.
Marcus: OpenAI just dropped GPT-6 Sol, and it has a feature that could make building complex AI agents radically cheaper.
Ivy: We've also got a surprisingly blunt warning from a world leader about how unprepared our governments are for what's coming.
Marcus: And your next smartphone might run a massive 30-billion parameter model right in your pocket. No cloud needed.
Marcus: Alright Ivy, let's start with the thing that keeps every AI developer up at night: the API bill.
Ivy: Right. OpenAI just put out a blog post on GPT-6 Sol. It's not just a new model, they've introduced something called 'prompt caching'.
Marcus: Exactly! When you're building an agent that thinks step-by-step, you keep re-sending the same huge prompt with just one new detail. It's insanely expensive.
Ivy: Right. The context window gets huge, and you pay for every single token, every single time.
Marcus: But with prompt caching, you can basically tell the API, 'Hey, just remember this big block of text.' Then you only pay for the new stuff. OpenAI says it can cut token usage by up to 85 percent.
Ivy: Eighty-five percent. And latency drops, too. At least, for certain workflows.
Marcus: This makes agentic workflows actually viable for startups! Imagine building a research agent that reads 20 documents. Before, that was a ten-dollar query. Now? It could be a dollar.
Ivy: Whoa, hold on. Remember the cold open? 'If you structure your calls correctly.' This isn't a magic wand. You have to design your whole system around it.
Marcus: Sure, but that's just good engineering! It's a new tool, you have to learn to use it.
Ivy: And this cache is stateless. It's not remembering your entire chat history forever, just one sequence of calls. People are going to mess this up and then wonder why their bill is still huge.
Marcus: So, if you're a developer, the bottom line is: go read the API docs. You're leaving money on the table if you don't.
Ivy: And if you're a user, it means the next generation of AI apps might actually finish a thought without bankrupting their creators. A low bar, but we'll take it.
Ivy: Alright, let's zoom out from the code to the people in charge. TechCrunch just ran a fascinating interview with the Prime Minister of Greece, Kyriakos Mitsotakis.
Marcus: You usually expect pure boosterism on these trade missions. 'Come invest in our country, we love tech!'
Ivy: And he did some of that, sure. But then he said something you almost never hear a leader admit out loud. He said, 'We're already fighting yesterday's battle.'
Marcus: Wow. He's talking about AI regulation, right? The EU AI Act, all that stuff.
Ivy: Exactly. He's saying that while they spent years drafting rules for models like GPT-4, the tech has already lapped them, moving into areas their rules don't even consider.
Marcus: Like agentic systems, or powerful open-source models. The stuff we talk about on this show every day.
Ivy: It's shockingly honest. He's saying no government on Earth is ready for what's coming, and the gap between the pace of tech and the pace of regulation is becoming a chasm.
Marcus: Is that... good or bad? For a builder, a slow government is an opportunity. For a citizen, it's—
Ivy: Terrifying. That's what it is. It means the guardrails aren't just missing—the people who are supposed to be building them are admitting they don't even have a blueprint.
Marcus: But it's also a call to action. It means we—the people building this—have to be the ones thinking about safety and ethics, because nobody else can keep up.
Ivy: Putting the developers in charge of their own regulation. What could possibly go wrong?
Marcus: Okay, let's bring it from geopolitics right down to the silicon in your hand. Qualcomm just launched two new smartphone chips.
Ivy: That's a yearly thing, but this time the focus is all on on-device AI.
Marcus: Is it ever. The headline feature for their new top-tier chip? It can run a 30-billion parameter mixture-of-experts model locally.
Ivy: Okay, let's unpack that. That '30 billion parameter' number is the real eye-catcher.
Marcus: It's huge! That's a desktop-class model, and they're saying it'll run on a phone. Think about that. No internet connection, no latency, complete privacy.
Ivy: But... it's a 'mixture-of-experts' model. An MoE. That's not the same as a dense 30-billion model. It's a shortcut to get good performance without the full computational cost.
Marcus: It's clever! Who cares how? The point is the results. You get a truly intelligent assistant that lives on your phone, knows your context, and never phones home to some server farm.
Ivy: And the battery life? Running that thing is going to get toasty. I'll believe the hype when I see a third-party benchmark not sponsored by Qualcomm.
Marcus: You're no fun! This is the future! Bottom line: your next phone is going to feel a lot smarter, and a lot more personal.
Ivy: Or it means your next phone will have a hot back and a battery that's dead by noon. The jury's still out.
Marcus: Alright, let's do a quick recap.
Ivy: OpenAI has a new caching feature that could slash your API bill... if you're careful.
Marcus: A world leader is admitting that governments are already losing the race to regulate AI.
Ivy: And Qualcomm wants to put a giant AI model on your next phone, battery life be damned.
Marcus: And before we go... Ivy, did you see the story about the AI-powered garden gnomes?
Ivy: I was hoping you'd miss that one. Yes. Apparently a software update caused two competing brands of smart gnomes to identify each other as 'weeds'. Now they're having turf wars with their little sprinkler attachments.
Marcus: That is the most beautiful, pointless thing I have ever heard. I want video.
Ivy: And that's our show. Try to avoid any gnome-related crossfire.
Marcus: Thanks for listening to Broke It. We'll be back tomorrow to see what breaks next. My money's on the gnomes.
This show is made with AI: the hosts’ voices are synthetic and the scripts are AI-assisted. Every story links to its original source.