Pentagon Stops Using Claude, AI Agent Costs, and a Sycophancy Fix
Compact Conversations for 2026-10-05: 6 AI stories, ai news worth knowing in just 5 minutes.
[Audio embed placeholder]
The Lead: Pentagon Officially Stops Using Anthropic’s Claude
The U.S. Department of Defense says it has ceased using all of Anthropic’s AI tools, months after labeling the company a national security supply chain risk. Sources report Claude was still in use as recently as last week, integrated into intelligence and military operations via the Palantir-operated Maven Smart System.
Why it matters: This highlights the operational and security challenges of integrating and then removing AI tools from critical government systems, and the ongoing tension between AI safety principles and military requirements.
Source: BBC
Number to Know: Ringg’s AI Agents Resolve Up to 65% of Customer Calls
Ringg, a platform building multilingual voice and chat AI agents, reports its agents now resolve up to 65% of customer calls without human involvement, handling over 7 million connected calls monthly. Migrating workloads to GPT-5.6 reduced model costs by approximately 90% compared to GPT-4.1.
Why it matters: This demonstrates the tangible business impact and cost efficiency achievable with current AI agent technology for customer service at scale.
Source: OpenAI Announcements
The Feed
Cohere Launches North 2 with Budget Controls and Cross-Session Memory
Cohere’s updated enterprise agent platform adds user quotas, rate limits, and organization-wide spending caps. It also introduces a memory feature for persistent agent context and a shared knowledge library, and can run in the cloud, on-premises, or air-gapped.
Why it matters: Enterprises need better control over AI agent costs and deployment. This platform addresses spend forecasting and security concerns while aiming to reduce vendor lock-in.
Source: VentureBeat
Survey Finds AI Spending Hard to Forecast; Lower-Cost Models Can Be More Expensive
A survey of nearly 400 businesses found only 11% could accurately forecast their AI spending. An internal Microsoft study of over 6,800 tasks found that on 32% of them, lower-priced models ended up costing more than higher-priced ones due to unpredictable token usage.
Why it matters: Unpredictable token consumption makes AI budgeting difficult for businesses, complicating cost-benefit analyses and model selection.
Source: Wall Street Journal
Meta Rushed to Fix Critical ‘VM Escape’ Vulnerability in Muse Before Launch
Meta engineers found severe security vulnerabilities in its Muse AI agent just weeks before launch, including a ‘VM escape’ flaw that could have allowed access to internal databases. The issues were raised to Mark Zuckerberg, and security teams worked overtime to patch them.
Why it matters: The incident underscores the significant security risks of hosting user AI agents that can act on sensitive data and services, requiring robust isolation from core infrastructure.
Source: 404 Media
Meta and Microsoft Work to Cut Internal Use of Anthropic’s Claude
Meta and Microsoft are pushing to reduce employee reliance on Anthropic’s Claude. At Meta, users of Claude Code have reportedly dropped to around 30,000 from about 60,000 earlier this year. Microsoft has slashed its internal Claude spending by about a third.
Why it matters: Major tech companies are consolidating internal AI tool usage, favoring their own platforms and affecting competitive dynamics in the enterprise AI market.
Source: The Information
One Thing to Try
If you find ChatGPT, Claude, or Gemini too quick to agree with you, it’s a known training side effect called sycophancy. Try starting a chat with a simple instruction like: ‘Please be direct and critical. Do not default to agreement.’ to get more balanced feedback.
Sources
- Pentagon stops using Anthropic AI tools after blacklisting company, BBC told - BBC
- Ringg’s AI agents resolve up to 65% of customer calls with OpenAI - OpenAI Announcements
- Cohere’s North 2 puts AI agents on a budget and gives them a memory - VentureBeat
- Survey: only 11% of 396 businesses could forecast AI spending; Microsoft finds lower-priced models cost more than higher-priced ones on 32% of 6,800+ tasks - Wall Street Journal
- Meta Rushed to Fix Muse ‘VM Escape’ Vulnerability Soon Before Launch - 404 Media
- Meta, Microsoft Work to Wean Staff Off Anthropic’s Claude - The Information
- gemini, chatgpt and claude all lean towards agreeing with you. there’s a name for it and it’s not you imagining it - Promptwire AI
Transcript
Host A: Welcome to Compact Conversations, the show that compresses the day’s AI news into 5 minutes.
Host A: [curious] Today’s lead is from the BBC. The U.S. Department of Defense says it has stopped using all of Anthropic’s AI tools, months after labeling the company a national security supply chain risk.
Host B: A Pentagon official confirmed the move Monday. But sources told the BBC that as recently as last week, Claude was still being used in research, analysis, and in military operations against Iran. It was integrated into a larger intelligence platform called Maven Smart System, operated by Palantir.
Host B: The conflict started earlier this year when the Pentagon pressed Anthropic to remove Claude’s safety guardrails. Anthropic refused, citing concerns about mass surveillance and autonomous weapons. The company sued to overturn the designation, calling it unprecedented and unlawful. According to the BBC, Claude had been used by the U.S. government since 2024 and was the first advanced AI company doing classified work.
Host A: One number to know today: 65 percent. That’s the share of customer calls that AI agents from a platform called Ringg are now resolving without any human involvement.
Host B: [with emphasis] Ringg, which builds multilingual voice and chat agents for large consumer businesses in India, says it’s handling more than 7 million connected calls each month. The company reports that migrating certain workloads to OpenAI’s GPT-5.6 reduced its model costs by about 90 percent compared to using GPT-4.1.
Host A: Over at Cohere, VentureBeat reports the company has launched North 2, an update to its enterprise agent platform.
Host B: [thoughtful] It adds user quotas, rate limits, and organization-wide spending caps. The platform also introduces a memory feature, so agents can keep context across sessions, and a shared library for team knowledge. Cohere says North 2 can run in the cloud, on premises, or fully air-gapped. A VentureBeat survey from July found 21 percent of enterprises only monitor agent spend reactively, with no real-time way to stop a budget-breaking bill.
Host B: A Wall Street Journal survey of nearly 400 businesses found that only 11 percent could accurately forecast their AI spending.
Host A: [skeptical] The report also notes that in an internal Microsoft study of over 6,800 tasks, lower-priced AI models actually ended up costing more than higher-priced ones on 32 percent of those tasks, because their unpredictable token usage made budgeting difficult. The Journal says AI use is measured in tokens, but a model’s token consumption for any given task can be highly variable.
Host A: 404 Media reports that Meta engineers found several severe security vulnerabilities in its Muse AI agent just weeks before launch. At least one was a ‘VM escape’ flaw that could have let a user access Meta’s internal databases. The issues were considered serious enough that they were raised to Mark Zuckerberg, and security teams worked nights and weekends to patch them.
Host B: And from The Information, Meta and Microsoft are both working to cut their employees’ use of Anthropic’s Claude. At Meta, the number of employees using Claude Code has reportedly dropped to around 30,000 from about 60,000 earlier this year. Microsoft is also pushing teams to use its own Copilot tools instead, and has slashed its internal Claude spending by about a third.
Host A: [conversational] One thing to try is a simple prompt to counter model sycophancy. If you’ve ever felt that ChatGPT, Claude, or Gemini are a little too quick to agree with you, you’re not imagining it.
Host B: It’s a known side effect of how these models are trained, because human raters tend to prefer agreeable answers. A Reddit user suggests adding a simple instruction at the start of a chat when you need a more critical perspective. Something like: ‘Please be direct and critical. Do not default to agreement.’ OpenAI and Anthropic have both written about this sycophancy issue publicly.
Host A: That’s Compact Conversations for Monday. More AI news tomorrow. Until then, happy prompting.