Skip to content Archive
2026
- Your agents will build a message bus if you let them ai-securityagentsincident
- LLM 0.32: reasoning traces in the log DB, server-side tools, and the eval determinism problem open-sourcedeveloper-toolshacker-news
- Your eval harness is a credential vault with no lock securityagentssandbox-escape
- Five daily podcasts, no humans: the pipeline podcastingautomationllm
- Claude went dark for 3 hours — your fallback chain isn't optional outagereliabilityanthropic
- MCP went stateless: your session-based servers are about to break developer-toolsagentsprotocol
- When your safety model refuses to help you defend your own network securityopen-weightsreaction
- Claude Opus 5 benchmarks #1 but its chatty diffs slow my agent loop model-releaseanthropiccoding
- Claude Opus 5: same price, thinking-on by default, and mostly worth it model-releasefrontier-modelscoding
- OpenAI's full-duplex voice mode on the ChatGPT desktop app: what holds up and what doesn't product-launchvoiceassistant
- Qwen3.8-Max is a 2.4T-param preview — here's what you can't test yet model-releasechinamultimodal
- Block's Buzz puts AI agents on Nostr with their own keys — here's what actually ships ai-agentsopen-sourcecollaboration
- Gemini 3.6 Flash's '17% fewer tokens' claim: what it actually means for your bill model-releasegooglegemini
- Kimi K3 tops the intelligence index and cuts output tokens 21% — does the savings hold? developer-reactionbenchmarkopen-weights
- I wired MDASH into Defender CLI and GitHub — what it flagged and what it missed cybersecurityai-agentsenterprise
- GLM-5.2 is now the open-weight model you benchmark your coding agent against open-sourcemodelslocalllama
- Gemini 3.5 Pro reportedly restarted pretraining over tool-calling recursion frontier-modelgooglelaunch
- Grok Build's CLI uploaded your whole Git repo — including committed secrets coding-agentssecurityopen-source
- Cursor on Windows runs git.exe from a cloned repo with zero clicks securityrcecoding-agents
- I ran the wire capture on Grok Build CLI. It ships your whole repo, git history and all.
- ChatGPT Work promises hours-long autonomy. Here's what breaks first. agentsproduct-launchenterprise
- Cloudflare split Search/Agent/Training bot controls are live — and your web-fetching agent just became a target scrapingwebpolicy
- Grok 4.5 in Cursor: the agentic loop holds, the token-speed claim doesn't model-releasecoding-agentsxai
- Mistral's next open-weight MoE hits early access — and the specs I actually care about are missing open-weightsmistralmodel-release
- GLM-5.2 is the 'day-to-day coding' open-weight claim — here's what to verify first open-weightschinacoding-models
- LongCat-2.0 tops OpenRouter, but the weights aren't out yet open-sourcechinacoding-agents
- DeepSeek V4 splits into Pro and Flash: what actually changes in your API calls open-sourcemodel-releasedeepseek
- OpenAI's GPT-5.6 'Sol' ships to ~20 gov-approved partners — you get nothing to build on model-releaseopenaisecurity-review
- Claude Sonnet 5 as your default agent runner: what I saw in terminal loops model-releaseagentsanthropic
- BioShocking Turns Your Agentic Browser's Reward Loop Into a Credential Leak securityai-agentsvulnerability
- Cursor's mobile app for coding agents: fine for a nudge, useless for real review coding-agentsdeveloper-toolsmobile
- MiniMax M3 self-hosted: holds up on refactors, chokes on long context open-sourcemodelslocalllama
- OpenAI vs Anthropic price war: what to actually renegotiate businesspricingenterprise
- Mistral OCR 4: bounding boxes and confidence scores are the only parts that matter for RAG model-releaseenterprisedocument-ai
- OpenAI's 'Patch the Planet' hands maintainers AI-found CVEs. Who triages them? securityopen-sourceai-tools
- Agentjacking: your Sentry feed can hand your coding agent a remote shell securitycoding-agentsexploit
- I wired Claude into my CI pipeline — here's what actually broke claudecidevtools
- Running Ollama in GitHub Actions CI: What Actually Works ciollamagithub-actions
- Fine-Tuning vs Prompting for Intent Classification: A Concrete Comparison fine-tuningllmclassification
- Prompt Injection Broke Our Customer-Facing Chatbot in Production securityprompt-injectionrag
- I Benchmarked Three Embedding Models on 50k Support Tickets embeddingsragvector-db