Ultrafast GPT, Anthropic’s Lead, and a Legal Prompt Injection

Compact Conversations for 2026-08-13: 7 AI stories, ai news worth knowing in just 5 minutes.

[Audio embed placeholder]

The Lead: OpenAI Launches GPT-5.6 Sol Ultrafast Mode: 14x Faster, 750 Tokens/Sec via Cerebras

OpenAI launched an Ultrafast mode for its flagship GPT-5.6 Sol model, using a new inference system from Cerebras to achieve 750 tokens per second—about 14 times faster than the standard version.

Why it matters: This preview is aimed at developers needing extremely low-latency responses for real-time agentic workflows, though it will come at a higher cost due to specialized hardware.

Source: OpenAI

Number to Know: Ramp’s July AI index: Anthropic’s market share hit 43.5%, widening its lead over OpenAI; Fable 5 is only 6% of tokens businesses bought

According to Ramp’s spend data, Anthropic’s market share among U.S. businesses reached 43.5% in July, widening its lead over OpenAI at 39.7%. However, Anthropic’s top-tier Fable 5 model made up only 6% of tokens businesses purchased.

Why it matters: The data suggests Fable 5’s high cost has hit a new upper bound for what businesses are willing to pay for AI performance, indicating price sensitivity even for the best available models.

Source: Ramp

The Feed

Google Launches Gemini 3.7 Flash With Major Coding Gains at Half the Price of 3.6 Flash

Google released Gemini 3.7 Flash, claiming major improvements in coding and reasoning at about half the price of the previous 3.6 Flash model for equivalent performance.

Why it matters: It’s positioned as a cost-effective option for developer and agent workloads where speed and price are prioritized over absolute top-tier capability.

Source: Google

Databricks closed a $5B funding round at a $190B valuation, six months after raising $5B at a $134B valuation, and says it has crossed $7B in revenue run rate

Databricks closed a $5 billion funding round at a $190 billion valuation, reporting it has crossed a $7 billion annual revenue run rate with over 80% year-over-year growth.

Why it matters: The funding fuels enterprise AI capabilities amid what the CEO calls ‘crazy’ demand for AI agents, highlighting the company’s Lakebase database and AI Gateway cost-control tool.

Source: CNBC

DeepSeek raises prices, adding dynamic pricing, ahead of a potential IPO; V4-Flash output tokens go from $0.28/1M to $1.32 during peak hours and $0.66 off-peak

DeepSeek is significantly raising prices for its flagship V4 models and introducing dynamic pricing ahead of a potential initial public offering.

Why it matters: The move, which sees output token costs increase multiple times, reflects market dynamics and cost pressures as AI providers seek to balance performance with profitability.

Source: Bloomberg

Microsoft begins merging its consumer and commercial Copilot apps into a single app, with a mobile and web rollout in mid-August and desktop in mid-September

Microsoft has started merging its consumer and business Copilot apps into one unified application, with a phased rollout through September.

Why it matters: This unification lays the groundwork for Microsoft’s upcoming ‘Super App’ and attempts to create a more cohesive AI product to better compete with ChatGPT, Gemini, and Claude.

Source: GeekWire

A person representing themselves in a Connecticut court hid prompt injections—tiny white text instructing any reviewing AI to side with them—within a legal filing.

Why it matters: The court caught the attempt, sanctioned the filer, and issued a warning about how such AI manipulation could become more common as AI tools integrate into legal workflows.

Source: 404 Media

One Thing to Try

When building reviewer agents, prompt them to act as a strict senior engineer tasked with finding deviations from the original spec, rather than using a generic review request. This small shift from a friendly peer to a designated critic can help break the ‘rubber-stamp’ pattern and catch real spec drift and over-engineering.

Sources

Transcript

Host A: Welcome to Compact Conversations, the show that compresses the day’s AI news into 5 minutes.

Host A: [curious] Today’s lead is a speed boost from OpenAI. The company launched an Ultrafast mode for its flagship GPT-5.6 Sol model, using a new inference system from Cerebras to hit 750 tokens per second—about 14 times faster than the standard version. OpenAI is framing this as a preview for developers who need extremely low-latency responses, especially for real-time agentic workflows.

Host B: It’s available now in the API as a separate configuration from the standard endpoint. OpenAI hasn’t detailed the pricing for Ultrafast yet, only that it will be higher than the base model because of the specialized hardware involved.

Host A: One number to know today is 43.5 percent. That’s Anthropic’s market share among U.S. businesses in July, according to new spend data from Ramp. [with emphasis] The analysis shows Anthropic widened its lead over OpenAI, which came in at 39.7 percent.

Host B: But here’s what caught the report’s attention: Anthropic’s top-tier Fable 5 model, released last month, made up only 6 percent of the tokens businesses actually purchased. At roughly ten dollars per million tokens, Ramp’s lead economist says the model has hit a new upper bound for what businesses are willing to pay, even for the best performance available.

Host A: Google launched Gemini 3.7 Flash with major gains in coding and reasoning. [thoughtful] The company claims it’s about half the price of the previous 3.6 Flash model for equivalent performance.

Host B: It’s available now through Vertex AI and AI Studio. Google is positioning it as a cost-effective option for developer and agent workloads where speed and price matter more than absolute top-tier capability.

Host A: Databricks just closed a five billion dollar funding round at a 190 billion dollar valuation—six months after its last five billion dollar raise at 134 billion. The company says it’s crossed a seven billion dollar annual revenue run rate, growing more than 80 percent year-over-year.

Host B: CEO Ali Ghodsi told CNBC the funding goes toward enterprise AI capabilities, citing what he called ‘crazy’ demand for AI agents. He also highlighted strength in the company’s Lakebase database unit and its AI Gateway tool for controlling model use and costs.

Host A: DeepSeek is raising prices significantly and adding dynamic pricing ahead of a potential IPO. For its V4-Flash model, output tokens are moving from 28 cents per million to a peak rate of one dollar and thirty-two cents during peak hours, with off-peak pricing at 66 cents.

Host B: Microsoft has started merging its consumer and commercial Copilot apps into one unified app. The rollout begins this week with Windows Insiders, then expands to mobile and web in mid-August, with desktop following in mid-September.

Host A: [lighter] And from 404 Media: a person representing themselves in a Connecticut court hid prompt injections in a legal filing—tiny white text telling any AI that might review it to side with them. The court caught it, the judge sanctioned the filer, and the ruling includes a thoughtful warning about how this kind of manipulation could become more common as AI tools spread through legal workflows.

Host B: One thing to try if you’re building reviewer agents is to give your reviewer a very specific, adversarial persona and its own copy of the original spec. Instead of a generic ‘review this code,’ prompt it to act as a strict senior engineer tasked with finding deviations from the requirements document.

Host A: [conversational] Several people in a Reddit thread about this exact problem reported that this shift—from a friendly peer to a designated critic—made their review loops actually catch real spec drift and over-engineering that the main agent produced. It’s a small prompt change, but it seems to break the rubber-stamp pattern.

Host A: That’s Compact Conversations for Thursday. More AI news tomorrow. Until then, happy prompting.