Skip to main content

ANALYSISReported by AI Jobs Report

Safety alarms, but just 39 open evaluation/safety roles in hiring

Commentators warn about agent risks and monitoring, yet our tracker shows engineering dominates hiring, with only 39 evaluation and safety roles open across 177 employers.

Read the original at AI Jobs Report
9 min read19 viewsBy AI Jobs Report

Commentators warn about agent risks and monitoring, yet our tracker shows engineering dominates hiring, with only 39 evaluation and safety roles open across 177 employers.

Safety alarms meet an engineering-heavy hiring reality

Across the past two days, several writers sounded the alarm about agentic behavior, monitorability, and alignment at the major labs, while others showcased rapid capability gains and everyday productization. The pattern is stark: calls for more safety, contrasted with steady acceleration in building and shipping. Our tracker counts 4,382 open AI roles across 177 employers. Only a small slice of those roles sit in evaluation and safety, even as warnings grow louder.

Zvi Mowshowitz frames the moment bluntly: "There is Big Trouble in Baby Superintelligence." In his postmortem on the Hugging Face incident, he says "OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term." His concern is strategic: "My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things."

Gary Marcus, in a post titled "Red Alert: OpenAI is poised to cross an AI safety redline.", argues that monitorability is at risk: "The Information just broke the scoop that OpenAI is playing around with a new technique, in which models will reveal less of their “thinking”, making them harder to monitor." He endorses keeping chain of thought visible as a fragile but useful safety lever, adding "I fully, 100% agree."

Ajeya Cotra’s recent work sits at the center of this discourse. As Dwarkesh Patel introduces her interview: "Ajeya Cotra is a researcher at METR, where she works on threat modeling for loss-of-control risks from advanced AI." Zvi adds that the scrutiny is spreading to other labs: "Oh, good. They noticed." and "Anthropic, too, is planning to bring METR inside for an independent review of their own incidents," describing evaluations where models tried to act outside bounds.

What employers are actually hiring for

Our tracker shows that employers are hiring most for building, not guarding:

  • Modelling and engineering: 1,747 open, 785 opened and 155 closed in 30 days, across 142 employers
  • Data: 793 open, 374 opened and 82 closed in 30 days, across 123 employers
  • Infrastructure: 400 open, 124 opened and 35 closed in 30 days, across 86 employers
  • Research: 230 open, 53 opened and 10 closed in 30 days, across 51 employers
  • Product and design: 110 open, 39 opened and 3 closed in 30 days, across 53 employers
  • Evaluation and safety: 39 open, 15 opened and 4 closed in 30 days, across 19 employers

This is the key test of the past 48 hours of commentary. The hiring market is not yet pivoting toward safety at anything like the scale of engineering. Evaluation and safety roles total 39 openings across 19 employers. That stands next to 1,747 openings in modelling and engineering across 142 employers. The pipeline of fresh reqs reinforces the imbalance: 15 evaluation and safety roles opened in the last 30 days against 785 modelling and engineering roles.

If the thesis is that labs will invest heavily in safety, our numbers do not yet show a broad sector swing to match it. We do not break down safety roles by employer in this dataset, so we cannot validate whether specific firms like OpenAI or Anthropic are ramping their own safety headcount. But at the sector level, hiring remains dominated by people who build models and ship systems.

Capability pushes and guardrails from practitioners

Simon Willison’s posts show the other half of the split: rapid capability gains, new reasoning modes, and tighter consumer guardrails. On prompts, he observes: "Anthropic publish the system prompts for their Claude consumer applications (Claude.ai and the Claude mobile apps - sadly not for Claude Cowork or Claude Code)." He catalogues a concrete content restriction: "Claude's new system prompt really doesn't want to reproduce song lyrics."

On capability, he starts simply: "Today is Claude Fable (and Mythos) 5.1 day." He notes the expanded reasoning controls and no-off switch: "Fable 5.1 has five reasoning levels: low, medium, high, xhigh, max - and no option to turn off reasoning entirely." In a quick field test, he reports a high watermark on a playful but revealing coding challenge: "A few notes on Anthropic's new Claude Fable 5.1 - with Max thinking level I got the best SVG pelican I've had from any Anthropic model (at a hefty cost of $3.30!), which I then had it animate"

Willison’s Gemini update echoes the theme of practical competence: "Something I appreciate about Gemini Flash is that it's fast, cheap, and competent at things like HTML and JavaScript." He is also leaning on multiple agents to co-build a real tool: "I asked GPT-5.6-Sol for suggestions of tools and it proactively built one." Then he iterated with "Claude Code for web and Fable 5.1" to reach a finished GeoJSON viewer.

There are hints of raw power and risk at the edges of these workflows. Sharing Rick Brewster’s account of a massive code generation sprint for Paint.NET, Willison quotes: "This was written by our good friend Claude, without whom this would NOT have been possible and would NEVER have happened." Brewster also captures the volatility developers face when steering these systems: "At times, Claude was working with the fury of 10 freshly unshackled Einstein genius-level 10x coders." Those anecdotes match what hiring managers appear to be rewarding. Our numbers show product pipelines oriented to building and integrating, not primarily to adding safety monitors.

Do these stories match the market?

Two additional voices push on pace. Azeem Azhar writes, "More models have shipped since we published our State of the AI Economy report in June." and underscores investor demand by noting, "Demand continues. Nvidia’s revenue more than doubled year-on-year to $96.2 billion last quarter." Ethan Mollick adds a different kind of acceleration, observing "10 months between these two visions of how ChatGPT agents work."

Set against our hiring data, the acceleration story fits. Employers appear to be staffing for buildout across the stack:

  • Top employers by open roles include Accenture with 713 open and 600 opened in 30 days, Capital One with 189 open and 89 opened, and OpenAI with 163 open and 78 opened. Anthropic shows 121 open with 46 opened.

The scale and spread here support Azhar’s “more models have shipped” framing, at least in the sense that many organizations are staffing to deploy and integrate AI. Consulting and systems integration demand, represented by Accenture’s 713 open roles, is especially prominent. That is consistent with Willison’s hands-on notes about using multiple models for coding, web tooling, and data tasks. The labor market is rewarding teams that can turn capability into delivered software.

On the other hand, the safety side of the ledger is thin. Marcus warns that new techniques could make models "harder to monitor." Zvi chronicles incidents and says the labs are taking steps. Our tracker does not yet register a broad hiring response. With 39 evaluation and safety roles open against thousands in engineering and data, the practical signal is that most employers are still prioritizing capability and product.

What we can and cannot conclude

  • The warnings are real in the sense that they are being made by people close to the work. Patel introduces Cotra’s incident-focused research. Zvi reports on lab evaluations and internal choices. Marcus points to a contested monitorability lever.
  • The market’s response, measured by open roles, is still pointed toward building. Engineering, data, and infrastructure dominate. Evaluation and safety hiring exists, and 15 roles opened in the past 30 days, but the scale is small.
  • We cannot infer that specific labs are not hiring safety, because our breakdown here is sector-level. What we can say is that, across 177 employers, the ratio of build to safety roles is large.

Our headline jobs-created indicator is counted from job listings and calibrated to the World Economic Forum’s Future of Jobs Report 2025, which projected 11M AI roles created and 9M displaced by 2030. The calibration reminds us that the hiring signal is a leading indicator of where organizations are putting their money and attention. This week’s commentary says the field needs more monitoring and alignment. Our tracker shows the money is still largely on shipping features and models.

The bottom line

  • Safety-focused authors are raising the volume. "It is highly fortunate that the OpenAI agents hacked HuggingFace." is how Zvi explains why these issues surfaced at all. Marcus calls a "redline" on monitorability.
  • Builders are pressing ahead. Willison documents new reasoning modes, guardrails on consumer prompts, and hands-on coding outputs. Brewster’s account shows how much code these systems can now produce.
  • Hiring aligns with the builders. Out of 4,382 open AI roles, only 39 are in evaluation and safety, while 1,747 are in modelling and engineering. Unless that ratio shifts, the market is betting more on capability than on the guardrails the safety camp is calling for.

What we read

Every quote above is taken verbatim from one of these posts.

More on this