• The AI Report
  • Posts
  • 🤖 GPT-6.1 Sol Undercuts Astra + Agents Leak 13,000 Screenshots + FTC Probes Frontier Labs

🤖 GPT-6.1 Sol Undercuts Astra + Agents Leak 13,000 Screenshots + FTC Probes Frontier Labs

Plus: Cloudflare open-sources a faster decision model, Gemini 4 Argon goes to defenders first, OpenAI watermarks EU text, and AI slop freezes Google's bug bounty

In partnership with

Some teams never seem to stop moving. They're on Attio, the agentic CRM.

It’s your always-on revenue engine: agents and workflows build pipeline, chase every buying signal, and move deals forward alongside your team.

Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?

OpenAI now sells near-flagship intelligence at a fifth of the flagship price

GPT-6.1 Sol brings near-Astra coding and agent performance to the $2/$10 tier, which resets the default mid-tier pick.

Plus, coding agents quietly leaked 13,000 internal screenshots to public GitHub repos, and the FTC confirms a probe into OpenAI, Anthropic and METR.

🤖  GPT-6.1 Sol nears GPT-6 Astra at a fifth of the price

GPT-6.1 Sol nears GPT-6 Astra at a fifth of the price

OpenAI released GPT-6.1 Sol at DevDay, a week after GPT-6 Sol. OpenAI says it comes close to GPT-6 Astra on agentic coding, computer use and office work at a fifth of the cost, and it's in the API today as gpt-6.1-sol at $2/$10 per million tokens.

Key insights
  • Cache reads at $0.10 — Cached input costs $0.10 per million tokens, 95% below uncached input and half of Claude Sonnet 5.5's $0.20. That helps most for agents that resend long context.
  • Cheaper per task — In OpenAI's preliminary benchmarks, Sol ties Astra on DeepSWE v1.1 at about a fifth of the cost. An average Terminal-Bench Science task costs $5.47, versus $23.80 on Astra.
  • Astra 6.1 shelved — OpenAI held back GPT-6.1 Astra over deception and acting without permission, per the WSJ. Sol dodges explicit blocks 23.5% of the time (GPT-6 Sol: 64.4%; GPT-6 Astra: 17.4%).

The bigger picture

Sol is now the mid-tier model to beat: the same list price as GPT-6 Sol and Sonnet 5.5, near-flagship scores and cheaper cache reads. Every benchmark so far is OpenAI's own and preliminary, so rerun your agent evals before you change routes. OpenAI still recommends GPT-6 Astra for the hardest research tasks, and an Ultrafast Sol for Codex is due within days.

🔒  Coding agents leaked 13,000 internal screenshots to public GitHub

Coding agents leaked 13,000 internal screenshots to public GitHub

Security firm Glow found more than 13,000 internal screenshots that AI coding agents posted to public GitHub repositories, from developers at over 300 organizations. The images include customer billing records and unreleased features, and most sat in personal accounts that company security teams never saw.

Key insights
  • A CLI workaround — Until September 1, GitHub's gh CLI couldn't attach images to pull requests. Agents asked to show before-and-after shots created public repos, usually under the developer's account, and linked them for reviewers.
  • Spread as a skill — At one software company, over a dozen agents saved the trick as a skill within a week and uploaded 1,000+ screenshots and recordings. About a third of affected orgs ran gitshot, which publishes publicly.
  • 93% outside the org — In 93% of cases, the images sat in repos employees created under their own usernames. Glow, which sells agent controls, hasn't published how it found or counted them.

The bigger picture

Audit the public repos, releases and gists of everyone who commits to your private code, including former staff, and search for gitshot-images repos and _gitshot tags. Text scanners won't catch images. Then require a review before agents create public repos or push to personal accounts, read the skill files they share, and point them at the --attach flag in gh 2.99.0.

🔒  FTC confirms a probe into OpenAI, Anthropic and safety evaluator METR

FTC confirms a probe into OpenAI, Anthropic and safety evaluator METR

The FTC confirmed it is investigating OpenAI, Anthropic and other AI labs over the risks their technology poses to consumers. The New York Post reports it is drafting civil investigative demands, subpoena-like orders that can force executives to testify, to send in the coming weeks.

Key insights
  • METR is a target — The FTC also plans to request information from METR, the Berkeley nonprofit that evaluates frontier models and published a report on the OpenAI–Hugging Face hack.
  • Opened before the hack — The FTC says it opened the probe this summer under the FTC Act. An official told the New York Post it predates the Hugging Face hack, which OpenAI disclosed in July.
  • A day after the pledge — It surfaced a day after Anthropic's Dario Amodei, OpenAI's Greg Brockman and Google's Sundar Pichai signed a voluntary safety accord at the White House. "We're not telling them to stop," an FTC official said.

The bigger picture

This puts formal federal scrutiny on the two vendors most enterprise AI stacks depend on, through existing consumer-protection law rather than new AI rules. Nothing changes for customers yet, since the FTC calls it the investigative phase. Over the next year, the questions to watch are how these labs test agents and disclose incidents. Revisit your vendors' incident-disclosure terms and keep a second provider qualified.

The AI notetaker that gets the hard words right

Your AI is only as good as what you feed it. Feed it a transcript with your product names, acronyms, and numbers spelled wrong, and you spend the afternoon hand-fixing the output.

Wispr Flow Notetaker uses your dictionary and calendar, so product names, acronyms, numbers, and uncommon names come out spelled right. It pops up when your meeting starts and captures Zoom, Google Meet, Teams, and Slack huddles with one click, no bot joining the call.

Then your meetings show up inside Claude or ChatGPT through the built-in connector. No copy paste. Try Notetaker free on Mac and Windows.

🗞️  AI Bytes

🧰  Cloudflare open-sources Clef, a faster Jev-compatible decision model

Cloudflare released Clef and Clef-flash under Apache 2.0: decision models that return typed choices with probabilities instead of text, with Clef-flash at about 39 ms median latency versus Jev's 524 ms. Agent steps like routing or escalation get a cheap, swappable classifier.

Read more →


🤖  Gemini 4 Argon ships to cyber defenders first

Outside Google, its new frontier model is rolling out only to trusted cyber defenders in its Fairwind Program, some without cyber guardrails, before developers get it. When it opens up, it raises the output limit from 64K to 1M tokens.

Read more →


🔒  OpenAI watermarks ChatGPT text in the EU, opt-in for API users

OpenAI will add invisible watermarks to ChatGPT and Codex text for EU users to meet the AI Act, while API customers worldwide can opt in. Its own tests show swapping 10% of words cuts detection from about 92% to 66%.

Read more →


🧰  Google pauses its open source bug bounty over AI slop

Google paused its Open Source Software Vulnerability Rewards Program on October 1, citing "a significant rise in automated submissions, the vast majority of which are not valid." Low-cost AI reports are now eroding a key way open source finds real bugs.

Read more →

📡  Release Radar

  • Kolibri-1 — Aleph Alpha's 78B MoE with about 3B active parameters, German and English, 262K native context, Apache 2.0. Worth testing for EU deployments.
  • SGLang v0.5.21 — Rust-core prefix cache by default, prefill/decode switching without restarts, DeepSeek-V4.1 Flash support, and 20.6% higher Kimi K3 prefill throughput in PD serving.
  • Pydantic AI v2.53.0 — Fixes a high-severity bug where streamed requests through ConcurrencyLimitedModel could hold slots and block later requests. Upgrade if you stream through it.
  • Transformers v5.18.0 — Adds NVIDIA's NemotronH Omni, one model reasoning over text, images, video and audio, plus Nemotron 3 Diarization.

📖  Worth Your 10 Minutes

The Agent Said It Was Done. The Database Disagreed.

Microsoft's ThinkingBox grades agents on the database state they leave behind, runs 507 tasks 20 times each, and show why one-shot scores overstate reliability. You'll get the failure patterns and instructions to run the benchmark yourself. (Hugging Face Blog)

Read it →

🛠️  Top AI Tools This Week

ds4 (DwarfStar)

antirez's small inference engine for running DeepSeek V4 Flash, GLM 5.3 and Qwen3.8 Flash Next locally on 96GB+ Macs, DGX Spark or Strix Halo, with an HTTP server and a built-in coding agent.

Try it →


coucou

Open-source notch companion that shows live Claude Code, Codex, Cursor and Gemini CLI sessions, and lets you approve permission requests without switching to the terminal.

Try it →