Skip to main content

ANALYSISReported by AI Jobs Report

Engineering hires dwarf safety as hype swirls around AI

Practitioners tout capability leaps, spending ceilings and trust crises; our 3,660-role tracker shows engineering demand surging, safety thin, and vendor hiring steady.

Read the original at AI Jobs Report
9 min read10 viewsBy AI Jobs Report

Practitioners tout capability leaps, spending ceilings and trust crises; our 3,660-role tracker shows engineering demand surging, safety thin, and vendor hiring steady.

A week of capability hype, spending ceilings, and a trust gap

Across the last two days, practitioners painted a picture of accelerating capability, contested business fundamentals, and a deepening trust problem. Simon Willison called the latest open model from Alibaba “a truly astonishing model,” while Gary Marcus dissected claims that Anthropic’s finances are surging, and Azeem Azhar argued enterprise spending is hitting its limits. Dario Amodei, quoted by Willison, warned that public skepticism is rooted in substance, not spin.

Here is the pulse of those claims, tested against our hiring tracker. We count 3,660 open AI roles across 142 employers, taken from their own job feeds. The pattern in hiring right now is concentrated engineering demand, paired with very modest investment in evaluation and safety.

Vendor hype vs hiring signals

Gary Marcus argues the narrative around one vendor is being pushed beyond what the evidence supports. He writes that “Anthropic is in an SEC-monitored “quiet period” before its IPO (expected sometime in the fall), somewhat limiting its communications,” but “that hasn’t stopped leakers and generative AI bulls from trying to convince audiences that Anthropic’s finances are spectacular.” He cites claims that Anthropic is “projecting 2028 ‌revenue of roughly $190 billion to $200 billion” and that “Gavin Baker claimed that Anthropic was making money on every token.” He also notes a widely shared projection that “Anthropic likely ends the year with ~$100-150B of revenue.”

Our hiring data cannot validate or falsify revenue claims. What it can show is whether headcount demand is moving. Among top employers in our tracker, OpenAI lists 152 open roles, with 54 opened in the last 30 days. Anthropic lists 104 open roles, with 19 opened in the last 30 days. Those figures indicate active, ongoing hiring at both companies but do not, on their own, support extraordinary revenue trajectories. They do rebut any notion that vendors have frozen hiring. They also suggest a gap between the intensity of the hype Marcus catalogs and what job postings alone can confirm.

Outside the model vendors, hiring is dominated by systems integrators and adopters. Accenture has 598 open roles (13 opened in the last 30 days). Capital One lists 179 (6 opened), Amgen 173 (3 opened), Waymo 106 (13 opened), Databricks 93 (17 opened), and PwC 87 (2 opened). If there is a single signal here, it is breadth of enterprise adoption rather than a single vendor’s breakout.

Are enterprises nearing a spend ceiling?

Azeem Azhar’s data note argues demand is skewed and budgets constrained at the high end. “The frontier races ahead. The top 10% of companies using OpenAI’s products use 8.3x more tokens than the typical firm.” At the same time, he writes, “Reaching the ceiling. Fable 5 token usage at businesses has been flat, making up only 6% of all their tokens (11% of spend) — businesses appear to have reached their limit on willingness to spend for the best model.”

Our tracker shows continuing investment in implementation talent despite any spend plateau on top-tier models. By role family, Modelling and engineering dominates with 1,391 open roles (1,558 opened and 167 closed in 30 days) across 115 employers. Infrastructure adds 333 open roles (379 opened and 46 closed) across 70 employers. Product and design, where you might expect cost-sensitive optimization and user-facing rethinks to show up, sits at 89 open roles (99 opened and 10 closed) across 47 employers.

Taken together, our figures are consistent with Azhar’s split: a small group may push token volumes hard, but across the broader market companies are still building teams to run, integrate, and tune models. Flat token spend on the most expensive frontier option can co-exist with rising hiring in engineering and infrastructure needed to operationalize a more mixed stack.

Open source, open weights, and who actually trains

Nathan Lambert draws a careful line between open models with full recipes and open-weight releases without training details. “The open-source language model – i.e. only models that come with a full training recipe, data, code, etc. – is a closer analogue to the open-source operating system. The open weight models you use – those with just model weights and inference code to run them – are closer to specific versions of software that you install in a project built upon them.” He argues recipes are costly but replicable: “The open-source recipe, typified in modern times by the Olmo models I helped build at Ai2, with its predecessors like Pythia from EleutherAI, is a resource intensive process that any company can pick up, modify, and press “run” on to produce a new set of model weights.”

Our hiring distribution supports the idea that many firms are implementers rather than primary trainers. Research roles total 202 open (218 opened and 16 closed in 30 days) across 44 employers. Data roles, essential for adaptation and evaluation, are 739 open (784 opened and 45 closed) across 97 employers. The far larger Modelling and engineering pool suggests more organizations are building products on top of available weights than investing in first-principles recipe training. Lambert’s note that “many companies are still using workflows built on Llama 3” aligns with a hiring pattern favoring application engineering over core pretraining research.

Trust, safety, and the data pipeline

Simon Willison highlighted Dario Amodei’s view that public skepticism is earned, not messaged away. Amodei wrote: “I think it is fundamentally a crisis of trust.” He added, “I think by far the most accurate criticism of AI companies including Anthropic is that we haven’t yet delivered on our big promises to benefit the world.”

Two adjacent posts underscore why that critique bites. Willison shared reporting that “Online forum discussions between Amazon workers confirmed that VGT3 destructively scans large volumes of books.” He introduces that investigation by noting, “We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility Excellent piece of reporting from 404 Media.” The appetite for training data is not theoretical; it is operational, and it invites scrutiny.

Our tracker shows companies are staffing the upstream pipeline far more than the downstream guardrails. Data roles account for 739 open positions across 97 employers. Evaluation and safety roles total just 40 open (45 opened and 5 closed in 30 days) across 16 employers. If restoring trust requires more measuring, red-teaming, and external validation, the market is not yet hiring for it at scale. That imbalance supports Amodei’s point about delivery and trust, and it shows up plainly in headcount demand.

Capability leaps, cost controls, and what that means for jobs

Simon Willison’s hands-on with Alibaba’s latest release captures the capability narrative: “Qwen 3.8 27B is a truly astonishing model.” Yet even there, operational realities intrude. On default settings, he reports, “The default of extra high results in spectacular over-thinking” and concludes “This is a hilarious default.” Those are cost and performance trade-offs that tend to get solved by engineering and infrastructure hires, not marketing. Our counts in those families are high. Product and design remains small by comparison, suggesting many teams are still getting the plumbing right before reimagining user experiences at scale.

Jack Clark’s research lens on discovery benchmarks points in the same direction: method and measurement work continues. He writes, “The new frontier for analyzing AI systems is understanding how good they are at inferring the unwritten rules of their environment” and asks, “How well can AI systems figure out the rules of their environment through exploration and curiosity, versus being fed them?” Our 202 open research roles show steady but not explosive demand for this kind of work inside employers’ pipelines.

What our data cannot confirm

Ethan Mollick observes that “AI diffusion in political campaigns is quite high,” and separately argues that “Google AI Overview alone would not profoundly change the nature of the web” is hard to imagine. Our tracker covers 142 employers’ own job feeds. It does not capture political campaign staffing in a way that lets us validate the first claim, and it cannot adjudicate long-run web dynamics from job postings. We can say that product and design roles are just 89 open across 47 employers, which does not yet look like a mass redesign of information products.

Emily M. Bender cautions against reifying “AI” as a singular thing: “I wish this weren't phrased as if "AI" were a thing.” That reminder applies here. Our numbers show specific hiring in specific families, not a monolith rising or falling.

Bottom line

This week’s commentary sets three tests. Are vendors’ business fundamentals as strong as boosters claim? Are enterprises pushing spend on the frontier, or settling into more sustainable stacks? And are companies investing in trust and safety in proportion to the risks they acknowledge?

Our tracker points to steady vendor hiring without evidence of extraordinary breakaway. It shows broad enterprise build-out in modelling, engineering, and infrastructure, consistent with Azhar’s picture of mixed spending behavior. And it shows a stark gap between data acquisition and evaluation and safety staffing. If trust will be rebuilt by delivery and measurement, as Amodei argues, the current hiring mix has catching up to do.

What we read

Every quote above is taken verbatim from one of these posts.

More on this