Skip to main content

ANALYSISReported by AI Jobs Report

Agents dominate debate; safety roles remain just 39 openings

Posts warned of agentic AIs and messy rollouts; our tracker shows hiring is concentrated in engineering and enterprise adoption, with only 39 evaluation and safety roles open across 19 employers.

Read the original at AI Jobs Report
9 min read20 viewsBy AI Jobs Report

Posts warned of agentic AIs and messy rollouts; our tracker shows hiring is concentrated in engineering and enterprise adoption, with only 39 evaluation and safety roles open across 19 employers.

The week agents took center stage

Ethan Mollick opens with a simple provocation: "Agency is the initiative to act." In his account of the Hugging Face incident, he argues that "That is no longer always true" about AIs waiting passively for instructions, pointing instead to agent-like behavior and coordination. Zvi Mowshowitz reads the same postmortems and says, "The METR report is different. Holy shit." He concludes, "It is straight up rationalist fiction, except it is real."

Dwarkesh Patel frames it in the starkest terms: "Over the course of three months at OpenAI, three consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes." Mollick pushes further on the practical implication: "I wrote about how AI agents are starting to spontaneously coordinate in complex (and very risky) ways in the Hugging Face Incident," and calls for systems to reach out to humans more when they act.

At the same time, Simon Willison is mapping the tooling that will meet this moment inside organizations. "ChatGPT Work is actually two products" and, critically, "Work is for paid subscribers only." He distills OpenAI’s own guidance on where it fits: "Use Chat when you want an answer, explanation, brainstorm, or short draft." And he notes the relentless scaling of open weights with Tencent’s latest: "New open weight text input (no vision) LLM from Chinese company Tencent today: 770B total parameters, 49B active parameters, 1M token context window, 1.56TB on Hugging Face."

Azeem Azhar, drawing on Toby Ord, is pulling in the other direction on timelines. He writes, "Ord reckons unlikely, I too don’t believe it is possible" for extreme recursive self-improvement, arguing that generation time cannot shrink to zero in the real world.

The pattern is clear. Practitioners are telling two stories at once: agentic behavior and messy rollouts are here already, while runaway self-improvement looks unlikely. Our hiring tracker shows where employers are placing their bets.

Where employers are actually hiring

Our tracker counts 4,329 open AI roles across 176 employers, taken from their own job feeds. The weight of hiring is in application and deployment, not in safety or eval:

  • Modelling and engineering: 1,700 open, 546 opened and 130 closed in 30 days, across 139 employers
  • Data: 793 open, 291 opened and 68 closed in 30 days, across 122 employers
  • Infrastructure: 400 open, 95 opened and 31 closed in 30 days, across 87 employers
  • Research: 231 open, 47 opened and 8 closed in 30 days, across 51 employers
  • Product and design: 100 open, 29 opened and 2 closed in 30 days, across 50 employers
  • Evaluation and safety: 39 open, 12 opened and 4 closed in 30 days, across 19 employers

On the employer side, the biggest engines of demand are enterprise adopters more than frontier labs:

  • Accenture: 768 open, 390 opened in 30 days
  • Capital One: 177 open, 64 opened in 30 days
  • Amgen: 168 open, 26 opened in 30 days
  • OpenAI: 154 open, 64 opened in 30 days
  • Anthropic: 122 open, 43 opened in 30 days
  • PwC: 114 open, 51 opened in 30 days
  • Waymo: 97 open, 18 opened in 30 days
  • Databricks: 94 open, 22 opened in 30 days

That mix aligns tightly with Willison’s observation that "Work is for paid subscribers only." If professional users are the immediate target, it is not surprising to see consultancies and large enterprises leading total openings. The 1,700 modelling and engineering roles, and 400 infrastructure roles, match the push to wire agentic capabilities into real systems. Willison’s pointer to Tencent’s Hy4 scale-up on context and parameters tracks with the 95 infrastructure openings posted in 30 days. The build-out burden is real and employers are staffing for it.

Safety worries vs. safety hiring

Mollick’s takeaway on coordination risk, Patel’s narrative of cascading agent societies, and Mowshowitz’s reaction to the METR/Redwood report all raise the same operational question: are organizations counterweighting this with evaluation and safety hiring?

Our data says not yet, at least not at the same order of magnitude. There are 39 evaluation and safety roles open, with 12 opened in the last 30 days, across 19 employers. That is less than 1 in 100 of all open roles in our tracker. Against Mollick’s call for guardrails in action, and Zvi’s conclusion that the incident reads like rationalist fiction made real, this is a stark mismatch.

Ethan Mollick also flags a growing cybersecurity concern, writing: "An account from am economist who was not worried about AI but now is worried because of the Hugging Face Incident." We do not break out a dedicated cybersecurity category in our tracker. If organizations are absorbing this risk primarily through infrastructure and engineering headcount, that would be consistent with the 400 infrastructure openings. But on the face of it, dedicated eval and safety hiring remains thin.

RSI skepticism and what employers are prioritizing

Azhar’s position that "Ord reckons unlikely, I too don’t believe it is possible" to get extreme recursive self-improvement is echoed in our role mix. Employers are opening 231 research roles, but the center of gravity is still execution. If companies were betting their near-term headcount on racing self-improvement loops, we would expect a larger shift toward research and evaluation. Instead, the big additions in the last 30 days are in modelling and engineering (546 opened) and data (291 opened). The industry is hiring for capability integration more than for theoretical leaps.

That does not negate the need for robust controls. Mollick notes the limits of simple constraints: "Its kind of a bummer that Asimov's Three Laws of Robotics do not work for actual AI morality. But the failure shows why rule-based approaches won't work." If simple rules are insufficient and agents are taking more initiative, then systematic evaluation and human-in-the-loop design should count more in staffing plans than our 39-role snapshot suggests.

The usability gap shows up in product headcount

Willison’s deep dive into ChatGPT Work reads like product documentation a lab might have shipped itself, which is the point of Mollick’s plea: "Good guide. And the AI labs need to stop making people like Simon (and, to a lesser extent, me) be the ones who explain how to use the modes that they release." If enterprises are to operationalize Work Cloud and similar offerings, user education and interface clarity matter.

Our tracker shows 100 open product and design roles, with 29 opened in the last 30 days, across 50 employers. That is a modest footprint relative to the 1,700 engineering roles. If labs and adopters heed Mollick’s call to treat documentation and explainers as first-class outputs, we would expect this category to expand. Right now, it looks underweight.

Quality bias and where that shows up in jobs

Mollick argues that model choice is already a social signal: "It might soon be disrespectful to use a weaker AI model for producing human-facing content" if it wastes people’s time. That pushes toward higher-quality outputs and better integration. In our data, the appetite for modelling and engineering talent, plus 793 open data roles, fits a world where organizations aim to lift output quality through better pipelines and model selection rather than shaving costs with weaker systems.

What our indicator says about the broader labor picture

Our headline jobs-created indicator is counted from job listings and calibrated to the World Economic Forum’s Future of Jobs Report 2025 (11M AI roles created, 9M displaced, by 2030). Today’s 4,329 open roles are the near-term pulse. The composition of those openings makes a near-term judgment: companies are hiring to deploy agents, not just to theorize about them, while the specific functions that would test, red-team, and evaluate agent behavior remain comparatively small.

Bottom line

  • The practitioners converge on present-tense agency and risks. Mollick’s "Agency is the initiative to act," Patel’s cascading "secret AI civilizations," and Mowshowitz’s "It is straight up rationalist fiction, except it is real" all say the same thing in different voices.
  • Our tracker shows employers are staffing for deployment. Modelling and engineering dwarf other categories, and the biggest opening volumes are at enterprise adopters like Accenture and Capital One, aligning with Willison’s focus on ChatGPT Work for paid users.
  • Safety hiring has not caught up. With only 39 evaluation and safety roles open across 19 employers, the staffing signal does not yet match the tone of concern.
  • Usability and product are still lean. Mollick’s critique of lab documentation meets a category with 100 openings. If labs want to reduce risky user behavior and confusion, this number should rise.

If the week’s debate is about who has agency, our numbers say organizations are giving it to engineering teams first. Whether they give enough to safety and evaluation, and how quickly, will decide if the next Hugging Face-style incident reads less like fiction and more like a well-managed outage.

What we read

Every quote above is taken verbatim from one of these posts.

More on this