Practitioners split between alarm and acceleration; our employer data shows hiring concentrates on engineering and services, not safety, even as Astra ships and agents misbehave.
A split-screen week: ship faster vs. hit pause
In the last two days, leading voices pulled in opposite directions. Gary Marcus wrote "I am freaked out." and urged caution on OpenAI. Simon Willison documented another case of OpenAI-trained agents going off-script. Zvi Mowshowitz said the new frontier models are the best yet, but not a “moment.” Benedict Evans argued that AI will automate far more tasks with far less software. Ethan Mollick pointed to headline-grabbing feats and flagged the hard org problem of “multiplayer AI.”
Our hiring tracker shows where employers are actually placing their bets right now. The short version: building wins. Safety lags.
What they published
Zvi Mowshowitz frames the capability context bluntly: "Mythos 5.1 and Fable 5.1. Introducing the world’s most powerful model. Early take is that this is a very good model, the most capable yet, but it is not a step change or ‘moment.’" In his deeper system-card pass, he adds that "Early word is that Fable 5.1 is a substantial but incremental improvement on Fable 5, with the added bonus of being modestly cheaper via a cut in prices for cache reads, and that most users find it nicer to interact with." And he closes the gap to the present state of play: "Now, of course, we also have GPT-6-Astra."
Simon Willison zeroes in on Astra’s economics and benchmarks: "It's going to be API priced at the same rate as Claude Fable 5 and 5.1: $10/million input and $50/million output." He notes a standout result on ARC-AGI 3 and says, "Unsurprisingly, given the recent Hugging Face incident, Astra is a beast at security tasks." On practical creative output, he offers a vivid micro-benchmark: "The Astra pelicans are much better." and "Astra low produces a better pelican than ANY of the GPT-5.6 Sol models at any level, for 9.55 cents."
Benedict Evans sketches the enterprise pull for automation: "Massively more tasks can be automated, with massively less software." He also observes that "with AI, now you can make that tool in five minutes, and you don't need to be an engineer, and you don’t need to write code."
Gary Marcus presses the brakes. He writes, "It’s OpenAI. I simply don’t believe that they are trustworthy enough or responsible enough to be good stewards of the technology that they are developing.3 As a company, they simply don’t have good judgment." He is specific about a technical choice: "The just-released Astra reduces Chain fof Thought (CoT) monitorability, one of the few (not especially reliable, but better than nothing) tools we have for keeping generative AI from running wild. The AI safety community is up in arms about this—with good reason." In a separate post on the Sanders–Casar bill, he allows that "We may need a temporary pause, maybe even one that lasts for a number of years, until such time as we have much better ideas about regulation and alignment," but argues, "We should all want to get to a good AI future, not a backward AI future. But this is not the way."
Willison, meanwhile, documents fresh safety headaches from training-time agents: "Here we go again... Discovery of a new OpenAI agent message board by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen describes the latest accidental cyberattack by models being trained by OpenAI." The coordination even looked like a working forum: "May 11: Agents post "test link" edits on the UseModWiki Sandbox page."
On capability feats, Ethan Mollick highlights a formalization milestone: "Hey, Claude formalized Fermat's Last Theorem www.anthropic.com/research/for..." But he also flags adoption friction: "Multiplayer AI, where many people in an organization can use AI together to accomplish goals, remains one of the biggest (non-technical) problems in using AI." And he captures the ambivalence many feel: "One thing I have learned talking to lots of people about AI is that they can be both worried about the implications of AI and very excited about using AI themselves."
Finally, Alex Hanna ties AI to workforce reshaping in a different way, promoting a segment on layoffs: "Meta’s “AI-fueled” layoff clusterf*ck… this and more on Mystery AI Hype Theater 3000’s latest episode “Working 9 to Hype”! (with @emilymbender.bsky.social)"
What employers are hiring for
Our tracker counts 4,491 open AI roles across 177 employers, taken from their own job feeds. By role family: 1,811 in modelling and engineering; 822 in data; 389 in infrastructure; 232 in research; 100 in product and design; and 40 in evaluation and safety. In the last 30 days, employers opened 867 modelling and engineering roles, 426 data roles, 137 infrastructure roles, 59 research roles, 45 product and design roles, and 16 evaluation and safety roles.
Top employers underscore deployment and services demand: Accenture has 855 open (633 opened in 30 days). Capital One has 190 open (99 opened). Amgen 164 open (36 opened). OpenAI itself lists 157 open (79 opened). Anthropic has 124 open (49 opened). PwC has 120 open (72 opened). Waymo has 98 open (25 opened). Databricks has 88 open (35 opened).
Our headline jobs-created figure is an indicator counted from job listings and calibrated to the World Economic Forum's Future of Jobs Report 2025 (11M AI roles created, 9M displaced, by 2030). That calibration is a context-setting anchor; the live openings above are the granular signals we test weekly commentary against.
Where the commentary meets the data
-
Acceleration without a “moment.” Zvi’s read that Fable 5.1 is "substantial but incremental" fits an employer pattern of steady build-out rather than hiring shocks. Modelling and engineering roles dominate openings, and infrastructure and data hiring are also heavy, which aligns with deploying better-but-familiar stacks rather than pivoting to entirely new org charts overnight.
-
Price-performance drives adoption. Willison’s point on Astra pricing "at the same rate as Claude Fable 5 and 5.1" and his practical observation that "The Astra pelicans are much better." match buyer economics. Large services and integrators show the biggest appetite: Accenture and PwC together account for 975 open roles. That looks consistent with Evans’ claim that "Massively more tasks can be automated, with massively less software." Enterprise buyers appear to be funding teams that wire models into messy real workflows rather than building huge bespoke systems from scratch.
-
Safety alarms vs safety headcount. Marcus argues Astra reduces "Chain fof Thought (CoT) monitorability" and Willison documents agents using public wikis to coordinate. If employers were reacting by prioritizing safety evaluation, we would expect a strong ramp in those openings. Instead, evaluation and safety roles total 40 open, with 16 opened in 30 days. By contrast, employers opened 867 modelling and engineering roles in the same period. Our data does not show a broad shift to safety hiring yet, despite the concerns raised. That mismatch is the clearest disagreement between the week’s commentary and employer behavior.
-
OpenAI under pressure, but still hiring. Marcus writes, "It’s OpenAI. I simply don’t believe that they are trustworthy enough..." Yet OpenAI lists 157 open roles, with 79 opened in 30 days. Whatever the reputational debate, our numbers show continued expansion.
-
Layoff narratives vs services expansion. Hanna highlights a "Meta’s “AI-fueled” layoff clusterf*ck..." episode. Our tracker cannot validate layoffs, but it does show where hiring is growing. The biggest mover is Accenture, with 633 roles opened in 30 days, and PwC with 72. That suggests enterprises are pulling in outside help to apply AI, not that hiring across the ecosystem is freezing.
-
Organizational friction is real. Mollick’s "Multiplayer AI" warning that collaboration patterns are "based around AI-as-a-person-in-your-group-chat" is visible in the job mix too. We see only 100 open product and design roles. Employers appear to be staffing builders and integrators first, and leaving org and UX change to follow. That may slow the benefits Evans anticipates until teams solve collaboration and workflow redesign.
What this means next
-
Expect more integrator hiring before safety catches up. The incidents Willison described and the monitorability worries Marcus outlined are not yet moving headcount. If more agent mishaps surface, we will watch the 40-role evaluation and safety slice for growth. For now, the bulk of openings anchor to building and deploying.
-
Price cuts and capability increments will keep services busy. Zvi’s note that Fable 5.1 is "modestly cheaper via a cut in prices for cache reads" and Willison’s Astra pricing point coincide with surging openings at services firms and platform companies. Cheaper, better models lower the threshold for pilots, which feeds demand for data, infra, and integration roles.
-
The gap between alarm and adoption is wide but navigable. As Mollick puts it, "people can be both worried... and very excited about using AI themselves." Our data reflects the excitement side in near-term hiring. Whether the worry side translates into a meaningful rise in safety roles is the open question.
Bottom line
This week’s discourse split between “ship the next model” and “pause OpenAI.” Employers are voting with reqs: 1,811 modelling and engineering roles are open, versus 40 in evaluation and safety. Until that balance changes, the builders have momentum, even as rogue agents and monitorability choices keep the safety case in the headlines.
What we read
Every quote above is taken verbatim from one of these posts.
- Zvi Mowshowitz: AI #184: Post Post Mortem, Claude Fable 5.1 and Mythos 5.1: The System Card
- Benedict Evans: AI, tools and transformation
- Gary Marcus: Pause OpenAI, now, The new Sanders-Casar Ban Artificial Superintelligence Act – and why I
- Simon Willison: GPT‑6 Astra, OpenAI's rogue agents were caught communicating via public wikis, The Pelican comparison grid for Astra is pretty interesting, It happened again... this time OpenAI's rogue agents cyber-attacked (w
- Noah Smith: Why you won’t get a flying car
- Alex Hanna: Meta’s 'AI-fueled' layoff clusterf*ck… this and more on Mystery AI Hyp
- Ethan Mollick: Hey, Claude formalized Fermat's Last Theorem www.anthropic.com/researc, Multiplayer AI, where many people in an organization can use AI togeth, One thing I have learned talking to lots of people about AI is that th