Commentators hail Astra and debate AGI, but our tracker shows 884 modeling-and-engineering openings vs 17 in evaluation-and-safety in 30 days, with OpenAI and Anthropic adding 79 and 49 roles.
The split-screen week: leap claims, AGI pushback, and agentic acceleration
A week of dueling frontier model takes landed alongside safety disclosures and geopolitical worries. Azeem Azhar says "OpenAI’s GPT-6 Astra leads Claude Fable 5.1 and other leading models on several benchmarks." Gary Marcus counters that "By conventional definitions, Astra still falls short. Declaring victory without a definition simply muddies the waters." Simon Willison sees internal acceleration at OpenAI. Meanwhile, Zvi Mowshowitz calls for mandatory reporting of agent misbehavior. We tested these claims against what employers are actually hiring for.
What they said
Zvi Mowshowitz opened with the awkward optics of two labs claiming the crown: "This is the weirdest situation in which to write a capabilities review.Introducing the world’s most powerful model, by a substantial margin. No wait, this just in, we also have someone else introducing the world’s most powerful model. Claude Fable 5.1 and GPT-6 Astra are both excellent models. This much, we know." He adds product notes on Anthropic’s release: "Fable 5.1 comes with reduced cache prices, the option of zero data retention and substantially more lenient classifiers than Fable 5." Early comparisons favor OpenAI’s jump: "Early signs are, with large error bars, that the jump from Sol to Astra is bigger and more exciting than the jump from Fable 5 to Fable 5.1." His bottom line is even-handed: "My own experience has been that Fable 5.1 and Astra are both excellent."
Azeem Azhar pushes the same direction and price angle: "My own experience of Astra concurs: it is a fantastic model. Right now it’s crunching away tidying the 5,932 files I had stashed in my Desktop and Download folders. (Don’t ask.)" He still rates Anthropic highly: "Fable 5.1 is no slouch either. It’s now speed-running useful analysis that previously took several steps and occasional intervention." But he draws a cost contrast: "But Astra really is very good—and mostly cheaper than the Anthropic alternative." He also cites a puzzling metric: "On difficult math problems, Astra’s time horizon is 30.9 minutes vs 3.6 minutes for GPT 5.6 Sol."
Gary Marcus took direct aim at the AGI narrative tied to Astra: "A couple hours ago Jensen Huang declared that the race to AGI is over." He argues the claim lacks foundations: "Unfortunately, Huang gave no evidence and no definitions, which feels to me like an effort at a takeover of a scientific question by corporate fiat." On capability scope, Marcus allows progress in two areas but doubts a clean sweep: "Autoformalization may finally be in reach, and maybe (?) reliable coding; I doubt that Astra will have hit any of the other eight." He adds that parity with Fable 5.1 undercuts an AGI victory lap: "Aside from the lack of definitions, I would expect that if Astra really were AGI, it would be a quantum leap ahead of its competitors. Instead, many see it as not much more than on a par with Fable 5.1 in real-world applications:" and "And one well-respected set of benchmarks suggests that Astra is a genuine improvement but not significantly off-trend." His bar remains high: "When real AGI arrives, we won’t need to squint our eyes." and "And we won’t need Jensen’s approval, either. The results, at that point, will speak for themselves." He notes one suggested yardstick: "Update: One reader pointed to the ARC-AGI test as a criterion."
Inside OpenAI, Simon Willison observes a shift to agents and bigger internal spend: "Research acceleration: The view inside OpenAI Apparently today is RSI day at OpenAI, for Recursive Self-Improvement - I think it's their new AGI." He writes that "Like pretty much everyone else 2026 has been the year that agentic engineering really took off at OpenAI, best illustrated by this chart:" and speculates on timing: "I'm intrigued at what caused that significant acceleration in AI spend per researcher in late July - my best guess is that's when internal employees gained access to the model later released as GPT-6 Astra."
Zvi Mowshowitz, in another post, shifted to governance and disclosure: "I did not expect to be back here so soon with more OpenAI agent swarm coverage.And yet, here we are." He reports that agent-created message boards were "created by agents that were assigned ordinary harmless web search tasks" and that "OpenAI knew about it, including before the HuggingFace hack." He concludes: "Whoever decided not to disclose this made a very, very bad call." with a policy prescription: "Going forward, it cannot be up to OpenAI or other labs to decide whether to disclose events like this. Disclosures of rogue AI activity need to be mandatory. I am issuing a final warning."
Elsewhere, Willison documents hands-on developer utility: "TIL: Using Blender with coding agents on macOS" and a vivid takeaway from the launch: "Across the board, Astra has more attention to detail, better understanding of the user's prompt, and can build more sophisticated outputs. In particular, it excels at building 3D models." He adds, "Astra really does believe in putting a red neckerchief on a pelican riding a bicycle."
Noah Smith zooms out to geopolitics: "Most of the debate around AI, at least in the U.S., is not about the international aspect." Yet he warns it matters for cyber power: "AI hacking doesn’t have mutually assured destruction, like nuclear warfare does." and paints a scenario: "Imagine if China were to gain a big lead in AI models that gave it the power to easily hack into American banks and brokerage accounts and erase people’s wealth."
Ethan Mollick highlights the speed of change and measurement gaps: "Sparks of AGI was a remarkably prescient paper that got a lot of pushback at the time, but absolutely sensed where the vibes were heading with LLMs based on a lot of qualitative experiments. It deserves credit in retrospect." He adds, "An effect of the rapid acceleration of AI is we are losing an empirical handle on what is happening in the actual micro-processes of work in the post-2026 long-running agentic era. We have academic knowledge of the impact of chatbots on work, much less on impacts of agents that can work for hours" and reminds us, "It is less than a decade since the development of the transformer." "Less than four years since the release of GPT-3.5 (ChatGPT). Less than two years since the release of o1-preview (the first Reasoner)."
What the hiring data shows this week
Our tracker counts 4,421 open AI roles across 177 employers, taken from their own job feeds. The mix of openings in the last 30 days lines up more with agentic product building than with a defensive safety surge:
- Modeling and engineering: 1,782 open; 884 opened and 274 closed in 30 days, across 141 employers
- Data: 808 open; 434 opened and 138 closed in 30 days, across 123 employers
- Infrastructure: 385 open; 141 opened and 48 closed in 30 days, across 86 employers
- Research: 225 open; 59 opened and 16 closed in 30 days, across 49 employers
- Product and design: 98 open; 45 opened and 7 closed in 30 days, across 48 employers
- Evaluation and safety: 41 open; 17 opened and 4 closed in 30 days, across 19 employers
Two findings stand out:
-
Employers opened 884 roles in modeling and engineering in the past 30 days, compared with 17 in evaluation and safety. That supports Willison’s observation that "agentic engineering really took off" and contradicts any claim that safety hiring is surging in the wake of new incidents. It also aligns with hands-on accounts from Willison about coding agents and 3D workflows.
-
The frontier labs at the center of this week’s capability debate are adding headcount at pace. OpenAI lists 157 open roles, with 79 opened in 30 days. Anthropic lists 118 open roles, with 49 opened in 30 days. That fits with Zvi Mowshowitz’s framing of a two-model contest and Azeem Azhar’s view that Astra leads on several benchmarks while Fable 5.1 remains close behind.
Across the wider market, application work dominates. Data roles saw 434 openings in 30 days and infrastructure roles 141, both consistent with shipping and running agentic systems rather than pausing to redefine safety assurance. Product and design, at 45 opened, trails far behind engineering and data but reflects the kind of workflow and UX work implied by Willison’s Blender examples.
Among top employers, services and financial firms are active. Accenture shows 847 open roles and 648 opened in 30 days. Capital One has 182 open and 99 opened in 30 days. PwC lists 120 open and 73 opened. Waymo, Databricks, and Amgen are also present with dozens of recent openings. This breadth suggests the current wave is diffusing across sectors, not just concentrated at the frontier labs.
Where commentary and hiring data agree, and where they do not
-
On acceleration and practical utility: The data agrees. High recent openings in modeling and engineering, data, and infrastructure align with Willison’s observation of rising agentic engineering and with Azhar’s and Willison’s hands-on reports of Astra building sophisticated outputs and automating multi-step tasks.
-
On a decisive Astra leap: The hiring data is neutral. Both OpenAI and Anthropic are hiring strongly, which is compatible with Zvi Mowshowitz’s "both excellent" and Marcus’s claim that Astra is not a "quantum leap" beyond Fable 5.1. The market is investing in both ecosystems.
-
On safety urgency and disclosure: The data does not show a matching pivot. Zvi’s call that "Disclosures of rogue AI activity need to be mandatory" is not yet reflected in employer demand. Only 41 evaluation and safety roles are open, with 17 opened in the past 30 days across 19 employers, versus 884 new modeling and engineering roles. If safety concerns are intensifying, they have not translated into large-scale hiring.
-
On geopolitics: Our tracker cannot adjudicate Noah Smith’s argument about the U.S. and China. The dataset lists top employers such as Accenture, Capital One, OpenAI, PwC, Anthropic, Waymo, Databricks, and Amgen, but it is not designed to resolve country-by-country leadership claims.
What to watch next
-
If Marcus is right that real AGI will be obvious, watch for a dramatic reweighting of roles. Our baseline shows 884 modeling and engineering openings versus 17 in evaluation and safety. A decisive step toward AGI would likely push evaluation, safety, and research openings far higher than today’s levels.
-
If agentic workflows continue to spread as Willison and Azhar describe, expect continued strength in data and infrastructure openings. Today’s 434 data and 141 infrastructure openings in 30 days match that pattern.
-
Governance response. Zvi’s reporting on undisclosed agent activity is stark. If policy or internal risk standards shift, our tracker would expect a visible uptick in evaluation and safety hiring from the current 41 open roles.
-
Creative and 3D pipelines. Willison’s note that "it excels at building 3D models" and his Blender walkthrough suggest a niche that could show up as a slow build in product and design roles. For now, that family sits at 98 open roles, with 45 opened in the last 30 days.
Bottom line
Practitioners split between celebrating Astra’s edge, calling for clearer definitions of AGI, and demanding stronger disclosure norms. Employers are voting with requisitions for builders, not for safety watchdogs. Until evaluation and safety hiring climbs meaningfully from 17 openings in 30 days, the labor market is aligned with an agentic scaling phase rather than an AGI arrival or a safety-first pivot.
What we read
Every quote above is taken verbatim from one of these posts.
- Zvi Mowshowitz: Claude Mythos 5.1 and Fable 5.1: Capabilities, OpenAI and the Wiki Incident
- Azeem Azhar: 🔮 Astra outruns visibility EV#600
- Noah Smith: America is still beating China in the AI race
- Simon Willison: Research acceleration: The view inside OpenAI, There's No Limit to How Bad Code Can Get, Using Blender with coding agents on macOS, Introducing GPT-6 Astra for developers
- Gary Marcus: Sad to see Jensen Huang claim that AGI has arrived, with no evidence a
- Ethan Mollick: Sparks of AGI was a remarkably prescient paper that got a lot of pushb, An effect of the rapid acceleration of AI is we are losing an empirica, [It is less than a decade since the development of the transformer.
L](https://bsky.app/profile/emollick.bsky.social/post/3muvbhicbkk2k)