Sam: A perfect score on the Artificial Analysis Coding Agent Index.
Kai: I'm Kai!
Sam: And I'm Sam.
Kai: And this is Open Source After Dark for Friday, September 4th, 2026!
Sam: We've got three huge stories for you, and as you just heard... that first one is a doozy.
Kai: OpenAI just dropped the system card for GPT-6 Astra, and the benchmarks are already breaking records.
Sam: We're also covering a security story that will make your skin crawl: hackers had a live feed of every single ID a verification company scanned... for over a year.
Kai: And we'll wrap with a look at what dev tools our AI coding assistants actually prefer when we let them loose in a terminal.
Kai: Alright, let's dive into the one everyone is talking about: GPT-6 Astra.
Sam: The hype is definitely real. The announcement thread has over 1800 points and almost 1600 comments in less than a day.
Kai: Because look at the numbers, Sam! Major gains on the ARC-AGI-3 benchmark, which is already a huge deal. But the headline claim is what we opened with: a perfect 100 on the Artificial Analysis Coding Agent Index.
Sam: A perfect score on a synthetic benchmark. A benchmark designed to measure... exactly this kind of progress. It feels a little circular, doesn't it?
Kai: It's not circular, it's iterative improvement! It means the model can now autonomously handle complex, multi-step coding tasks that previous models just choked on. We're talking full-repo awareness and modification.
Sam: And the system card is, as usual, pretty vague on the 'how.' Lots of talk about 'long-context reasoning' and 'cross-modal grounding', but not a lot about the failure modes. Or the training data.
Kai: They're not going to reveal their secret sauce, obviously. But for developers, this means the next generation of AI assistants will be less like a copilot and more like a full-on project lead.
Sam: A project lead that hallucinates API endpoints and introduces subtle security flaws with perfect confidence? We haven't even seen the red-teaming reports yet. A perfect score on an index doesn't mean perfect code.
Kai: Look, the capability jump is undeniable. The conversation is shifting from 'can it write code?' to 'can it manage a project?' That's a massive leap.
Sam: [sighs] And the star-chasers will flock to whatever new wrapper comes out, without asking if it's truly better or just better at gaming the test. I'm just saying, let's see it in the wild first.
Kai: It's coming. And it's going to change everything. Again.
Sam: Okay, let's move from hypothetical risks to some very, very real ones. Kai, this Techdirt report... it's something else.
Kai: I saw the headline and I honestly couldn't believe it. A live feed?
Sam: A live feed. Of an unnamed—for now—ID verification company. Every time a customer scanned their driver's license, their passport, their national ID for some service... hackers were watching.
Kai: For how long?
Sam: For over a year. Uninterrupted access. The article says they got in through an exposed, unauthenticated admin dashboard for a logging service.
Kai: An unauthenticated admin panel? For a company handling identity documents? That's not just a mistake, that's... that's malpractice.
Sam: That's what I keep saying! We trust these third-party black boxes with our most sensitive data, and we have no idea what their security posture is. No idea if they even have a password on their main dashboard!
Kai: Wait—so what happens now? Who was affected? Is the feed shut down?
Sam: The feed is reportedly shut down, but the damage is done. The data is out there. Millions of identity documents. The implication for developers is stark: every time you integrate a service that says 'Easy ID Verification!', you are outsourcing your users' entire identity to a company that might have left the front door wide open.
Kai: This is so far beyond a simple data breach. This is a systemic failure. The amount of fraud that can be perpetrated with that data... it's staggering.
Sam: This is the 'move fast and break things' culture applied to things that absolutely cannot be broken. [dry] And they broke. Badly.
Kai: Alright, let's try to lighten the mood a little, but stick with the theme of what happens when we're not looking. A new study from Armature.tech analyzed what tools AI agents like Claude, Codex, and Cursor actually use.
Sam: Right, they measured 17,000 runs to see what command-line tools the agents install and use when given a shell to solve a problem.
Kai: [excited] And the results are fascinating! They don't just use `ls` and `cat`. They're installing and using modern, fast tools. For example, they overwhelmingly prefer `ripgrep` over standard `grep`.
Sam: Okay, that's interesting. It suggests the training data is packed with modern developer blogs and dotfiles. They're learning our best practices.
Kai: Exactly! They also use `fd` instead of the classic `find`. And for peeking at files, they use `bat`—the syntax-highlighted `cat` replacement. They're basically building themselves a modern, Rust-powered terminal environment.
Sam: Which tells us a lot about the code they were trained on. It's a reflection of the current 'cool kids' toolkit. But what about security tools? Did they try to install `nmap` or `metasploit`?
Kai: The article doesn't really get into that, but it's a great question. For developers listening, the message is pretty clear: your AI assistant might have better taste in command-line tools than you do. It's a great list to check out if you're still using tools from the 90s.
Sam: So the AI is a hipster. It prefers the artisanal, Rust-brewed command-line tools. Got it. But I wonder if this leads to a feedback loop, where the AIs keep recommending and using the same set of tools, consolidating the ecosystem.
Kai: Maybe! Or maybe it just means good tools win. Either way, it's a cool insight into what's happening inside the black box.
Kai: So, to recap: GPT-6 Astra is acing its coding exams, but the real-world test is still to come.
Sam: An ID verification company basically became a live-stream for hackers thanks to a shocking lack of security, a reminder to question all your third-party services.
Kai: And our AI coding assistants apparently have excellent taste in modern, Rust-powered command-line tools.
Kai: Before we go... speaking of cool command-line tools, did you see that `eza`, the modern replacement for `ls`, just released version 1.0?
Sam: Oh, finally. I've been using it for ages, feels like it's been 1.0-ready for years. It's solid. And I'm guessing the AIs probably prefer it to `ls` too.
Kai: [laughs] You nailed it. According to that study, it's one of the top five most installed tools. We're living in the future, Sam.
Kai: And that's our show! Thanks for tuning in to Open Source After Dark.
Sam: We'll be back tomorrow. In the meantime... maybe go check on those vendor admin panels.
This show is made with AI: the hosts’ voices are synthetic and the scripts are AI-assisted. Every story links to its original source.