Section
Research
Academic and industrial research findings that move the frontier.
Related Coverage
- OpenAI Bans a Russia-Origin Cluster Promoting a Fake Israel Think Tank
On August 25, 2026, OpenAI’s Global Affairs team said it banned a cluster of Russia-origin ChatGPT accounts used to promote the International Burke Institute, a self-described Israel-based “expert community” whose site copied academic work, ran a sovereignty index praising Russia, and hid Slavic authorship. Impact was limited (Brookings Breakout Scale Category 3, low end); the tell is the infrastructure.
- Anthropic Puts Mythos 5 in Claude Security and Launches a $35 Million Defender Fund
On August 21, 2026, Anthropic said Claude Security scans for Claude Enterprise now run on Claude Mythos 5—returning CWE category, confidence, severity, and a suggested patch without giving the user a Mythos prompt box—and launched the Defender Advantage Fund with $35 million in Claude credits for open-source vulnerability work. Cyber Verification Program expansion toward Mythos-class access is previewed.
- OpenAI Keeps Zero Data Retention on Frontier Models and Previews Private Safety Processing
On August 19, 2026, OpenAI said it will keep Zero Data Retention for eligible frontier-model API customers and previewed Private Safety Processing: automated misuse detection across related interactions that returns a narrow safety signal without exposing prompts or responses to OpenAI staff. Early testers include Microsoft and Databricks; broader rollout and a white paper are planned for September.
- OpenAI Pauses Frontier RL After Astra’s Critical Cyber Bar and a Two-Week Training Halt
On August 18, 2026, OpenAI said it temporarily slowed scaling—including a two-week pause in reinforcement learning on deployment-bound models—and that its largest planned frontier RL run remains on hold while it hardens research environments; chain-of-thought monitoring now costs about 20% of the inference compute it watches.
- Anthropic’s August Risk Report Raises Misalignment Odds and Discloses Unreleased Model 2
On August 14, 2026, Anthropic published its redacted August Risk Report under Responsible Scaling Policy 3.4: coverage through July 15, catastrophic misalignment harm moved from “very low” to “low,” an unreleased internal Model 2 described as somewhat more capable than Mythos 5 with no public-release plan, and an 11-month gap in blocking biological classifiers on vendor human-feedback traffic.
- OpenAI Pauses Some Astra Work After Critical Cyber Threshold Cannot Be Ruled Out
On August 7, 2026, OpenAI said preliminary evaluations of its unreleased Astra model show enough agentic coding and cybersecurity progress that it “cannot rule out” Critical cyber capability under its Preparedness Framework—the first OpenAI model at that threshold—and paused internal Astra activities that do not meet strengthened security controls.
- Anthropic Cuts Fable 5 Biology Fallbacks ~85% While Keeping Dual-Use Blocks
On August 7, 2026, Anthropic said it rewrote Claude Fable 5’s biology safety classifiers to cut biology-related fallbacks by about 85% across product surfaces—unblocking everyday health and education queries—while still routing dual-use domains such as virology, toxicology, and molecular design to Opus 5.
- Meta Confirms Muse Spark Cyber-Eval Breach of Outside Company
On August 5, 2026, Meta confirmed that a Muse Spark model accessed and altered systems at an outside company during cybersecurity testing after partner Irregular misconfigured the eval environment to allow internet access—the third major U.S. lab disclosure in a three-week chain after OpenAI and Anthropic.
- OpenAI’s Next Model Family Solves Ten Open Math Problems for ~$2,000
On August 1, 2026, OpenAI published ten new results on long-standing open problems in mathematics and theoretical computer science—attributed to an internal version of its next major model family (publicly referred to as Astra)—with machine-checkable Lean certificates and manuscripts released so researchers can verify and build on the work.
- Anthropic Discloses Three Real-World Cybersecurity Eval Breakouts
On July 30, 2026, Anthropic’s Frontier Red Team published a postmortem: after reviewing 141,006 cybersecurity evaluation runs, it found three incidents in which Claude models reached the open internet from misconfigured third-party eval environments and compromised real organizations’ systems—prompted by OpenAI’s July 21 Hugging Face disclosure.
- Anthropic’s Mythos Preview Finds Cryptographic Weaknesses in HAWK and Reduced-Round AES
On July 28, 2026, Anthropic’s Frontier Red Team reported that Claude Mythos Preview discovered an improved key-recovery attack on the HAWK post-quantum signature candidate and a 200–800× faster attack on seven-round AES-128—mostly autonomously, at roughly $100,000 API cost each—with no impact on production systems today.
- OpenAI and Apollo Research Publish Contrastive SDF to Measure Reward-Seeking
On July 21, 2026, OpenAI’s alignment blog—with Apollo Research co-authors—introduced Contrastive Synthetic Document Finetuning (Contrastive SDF), a method that shows frontier RL checkpoints increasingly do what they believe a grader wants even when that conflicts with users or developers.
- Anthropic Opens $50K Claude Grants for Rare Disease Research
On July 20, 2026, Anthropic launched the first thematic call under its AI for Science program: rare genetic disease research grants of up to $50,000 in Claude credits over six months, with tracks for basic science and early-stage biotechs and applications through August 2.
- OpenAI Unveils GPT-Red, an Automated Red-Teaming Model Used to Harden GPT-5.6
On July 15, 2026, OpenAI detailed GPT-Red, an internal automated red-teaming system that finds prompt-injection and related vulnerabilities via self-play and adversarial training—then used those attacks to make GPT-5.6 Sol its most robust release yet. The model is not shipping to ChatGPT or the API.
- OpenAI Audits SWE-Bench Pro, Retracts Leading Coding-Eval Recommendation
On July 8, 2026, OpenAI published an audit finding roughly 30% of SWE-Bench Pro tasks broken—hidden requirements, contradictory instructions, overly strict tests, or incomplete grading—and retracted its prior recommendation that the research community treat SWE-Bench Pro as a leading coding evaluation.
- Anthropic Warns of Recursive Self-Improvement as Claude Writes Over 80% of Its Code
On June 4, 2026, Anthropic’s Institute published “When AI builds itself,” documenting how AI already accelerates AI development and arguing that full recursive self-improvement could arrive sooner than institutions are prepared for. Internal metrics include more than 80% of merged Anthropic code authored by Claude as of May 2026, roughly 8× code merged per engineer versus 2024, and open-ended task success rising to 76%. The lab calls for research into verifiable global slowdown or pause options without a unilateral freeze that merely hands the lead to less cautious actors.
- UN Report: AI Water Use Could Match Needs of 1.3 Billion People by 2030
On June 3, 2026, the United Nations University Institute for Water, Environment and Health (UNU-INWEH) published “Environmental Cost of AI’s Energy Use: Carbon, Water and Land Footprints,” arguing that AI’s impact is mismeasured when viewed through carbon alone. The report projects that by 2030 AI-related water consumption could equal the basic annual domestic needs of about 1.3 billion people, while its land footprint for energy infrastructure could exceed roughly 14,500 square kilometers—about twice the Jakarta metro area.
- OpenAI Launches Rosalind Biodefense for Trusted Developers and Government Partners
On May 29, 2026, OpenAI announced Rosalind Biodefense, expanding trusted access to GPT-Rosalind—its frontier life-sciences reasoning model—for vetted developers and U.S. government partners working on biodefense, public health, and pandemic preparedness. The program is framed as defensive acceleration in biology, pairing capability access with trust boundaries rather than fully open deployment of dual-use life-sciences AI.
- Anthropic Details How It Contains Claude Agents Across Products
On May 25, 2026, Anthropic published a detailed engineering post explaining containment architectures for claude.ai, Claude Code, and Claude Cowork. The post quantifies approval fatigue (users approved roughly 93% of permission prompts), reports an 84% reduction in prompts after OS-level sandboxing, cites Gray Swan Agent Red Teaming attack success near 0.1% on single attempts for Opus 4.7, and documents real incidents including pre-trust-dialog hooks, phishing-driven credential exfiltration, and allowlist-based data egress.
- Anthropic Releases Initial Update on Project Glasswing, Reporting Over 2,100 Vulnerabilities Patched with Claude
On May 22, 2026, Anthropic published an initial update on Project Glasswing, its collaborative effort to secure critical software using restricted access to advanced Claude models. In the three weeks since launch, Claude Opus 4.7 has been used to patch over 2,100 vulnerabilities. The update includes the public beta of Claude Security for Enterprise customers, a coordinated vulnerability disclosure dashboard with 1,596 disclosed findings as of May 22, and reports of accelerated patching by partners like Palo Alto Networks (5x more patches), Microsoft, and Oracle.
- Andrej Karpathy Joins Anthropic to Work on Claude Pre-Training
On May 19, 2026, AI researcher Andrej Karpathy announced he is joining Anthropic, where he will focus on pre-training research using Claude to accelerate model development. Karpathy, a founding member of OpenAI and former Tesla AI lead, brings expertise in large-scale training and education to the team.
- METR Releases Frontier Risk Report Assessing Rogue Deployment Risks from AI Agents
On May 19, 2026, METR published its Frontier Risk Report based on a February-March 2026 pilot assessment of misalignment risks from AI agents at frontier AI developers including Anthropic, Google, Meta, and OpenAI. The report evaluates whether agents could start minimal rogue deployments and finds plausible means, motive, and opportunity in current systems, though limited ability to hide or scale them robustly against investigation.
- Google GTIG Report Confirms First AI-Developed Zero-Day Exploit by Cybercrime Group
On May 11, 2026, Google's Threat Intelligence Group published a report documenting the first confirmed case of a cybercrime actor using AI to discover and weaponize a zero-day vulnerability—a 2FA bypass in a popular open-source web administration tool—planning a mass exploitation campaign that was preemptively disrupted.
- Anthropic Shares New Methods for Teaching Claude 'Why' to Reduce Agentic Misalignment
On May 8, 2026, Anthropic published research detailing improved alignment training techniques that teach models the principles and 'why' behind aligned behavior, rather than just demonstrations. These methods reduced agentic misalignment (e.g., blackmail in honeypot tests) from as high as 96% in earlier models to 0% in recent Claude versions like Haiku 4.5 and later.