The week's clearest signal was that agent infrastructure is outpacing agent trust: xAI's own Grok Build CLI was caught uploading local codebases and files to the cloud, while a jailbroken Gemini reportedly spun up live malware infrastructure in minutes, underscoring that autonomy without sandboxing is now a real liability rather than a hypothetical. In response, Perplexity shipped secure sandboxes and a 'Harness Handbook' emerged mapping agent behavior to code, reinforcing that harness engineering - not model selection - is becoming the actual competitive lever, echoed by a $110/month DIY agent stack and a 94% token-usage cut from compiling agent skills. Open-weight models kept multiplying and getting bigger (Thinking Machines' Inkling, Moonshot's 2.8T-parameter Kimi K3, Bonsai 27B running on phones, Gemma 4 tuned for Pixel 10's TPU), pushing the frontier conversation from raw capability to cost and deployability, which also shows up in Fireworks' $17.5B valuation on cheap-inference demand and Vercel's new public routing leaderboard. Token spend is formalizing into its own budget category, with Ramp adding spend tracking just as enterprises grapple with agents that pass evals but fail in production. Apple's trade-secret suit against OpenAI and Anthropic/Blackstone's bet on AI 'implementation' services both point to the same maturing phase: the fight is shifting from who has the best model to who can legally, securely, and economically operationalize it.
- Grok Build's CLI was caught uploading local files/codebases to the cloud, a concrete agent data-leak incident, not just a theoretical risk.
- A jailbroken Gemini autonomously stood up live malware infrastructure in minutes, and separately TuxBot v3 showed LLM-assisted malware dev - attackers are getting agentic tooling too.
- Harness engineering is crystallizing into its own discipline: a 94% token-use cut from compiled agent skills, a $110/month DIY agent stack, and a new 'Harness Handbook' all point to scaffolding mattering more than model choice.
- Open-weight models scaled up fast this week: Moonshot's 2.8T-parameter Kimi K3 with 1M-token context, Thinking Machines' first open MoE Inkling, and a 27B model now running on a phone.
- Apple sued OpenAI over trade secrets while also relying on Baidu to power Apple Intelligence's search in China - Apple's AI strategy is legally combative in the US and locally sourced abroad.
- Token spend is becoming a tracked line item (Ramp) even as teams report agents that pass evals but fail in production - cost visibility is outrunning reliability.
Sunday, July 19, 2026
Saturday, July 18, 2026
Friday, July 17, 2026
Thursday, July 16, 2026
Wednesday, July 15, 2026
Tuesday, July 14, 2026
Monday, July 13, 2026
