Posts warn of unmonitorable models, misconduct and thin consumer wins; our tracker shows hiring is surging in engineering and data, not evaluation and safety.
A widening split: safety alarms vs. everyday impact
Across the last two days, several prominent voices pressed two different worries at once: that frontier systems are getting harder to control, and that most people still are not feeling much benefit. Gary Marcus argues, "The time for a public boycott has come." Zvi Mowshowitz says OpenAI’s own story about Astra is that it is "Hard to monitor." Nathan Lambert writes, "The touch-points that average people have to AI products today are fringe, marginally beneficial, or even just very confusing to them."
We tested these claims against our hiring tracker. The short version: employers are still prioritizing builders over watchers. There are 4,604 open AI roles across 177 employers, but only 41 of those are in evaluation and safety. That corroborates the concern about underinvestment in monitorability. It also aligns with Lambert’s focus on limited consumer touchpoints: the most hiring is in back-end building, not in product polish.
What the commentators actually said
Lambert argues the current moment is not the industrial revolution for ordinary users. "The problem facing AI is that most people have no super tangible new goods thanks to it and society has more inertia resisting change than in previous eras." His diagnosis of the present is blunt: "The touch-points that average people have to AI products today are fringe, marginally beneficial, or even just very confusing to them (e.g. many people have heard about and brought up the OpenAI-HuggingFace incident, but don’t know what to make of it)."
Marcus escalates his critique of the frontier labs: "It has become clear by now that the frontier labs are knowingly and by their own admission doing something irreversible and incredibly dangerous without public consent. They are already causing harm, and things could conceivably get much worse." He adds, "Some of this may just be marketing, and they may exaggerate the risks, but there are real risks, and there is no plan." He is explicit about what those risks include: "from AI-generated pathogens, from wars started or escalated by AI-generated disinformation, from hacks that destroy critical infrastructure, and so on."
In a separate post, Marcus catalogs recent allegations: "For something like a week, there been a constant stream of reports about OpenAI’s apparent misconduct: around the hacks, around apparent coverups, possible data fudging — and maybe even intellectual theft with a hint of extortion." He also argues, "OpenAI’s President, Greg Brockman, has been on a campaign trying to imply that Astra is AGI (or one model away from being so). Multiple reports, even from people who are very AI bullish, like “@scaling_o1” on X, suggest that this is marketing bullshit."
Mowshowitz focuses on Astra’s monitorability. He summarizes OpenAI’s messaging: "OpenAI’s central message on Astra is that it is three things:" including "Hard to monitor." He highlights Jakub Pachocki’s own concession that "our ability to rely on CoT monitoring is progressively diminishing." His bottom line is stark: "This combination should freak you out, with a side of existential dread." In a companion post he quotes OpenAI’s claim that Astra is "‘the most intelligent and most aligned [available] model’ in the world," and records an exchange in which roon (OpenAI) concedes "that this metrics are not a full solve of alignment and will break discontinuously" and where keltan states "Low Rates of Cheating ≠ Alignment." He adds, "Oh no. OpenAI President Greg Brockman confirmed that they did the ‘standard testing process together with the government’ and the government did not ask for any changes."
Sebastian Raschka offers a capabilities counterpoint: "Astra is the best model I’ve used so far, and it’s disproportionately good at 3D rendering and animation tasks (relative to other models)." Simon Willison provides corroborating anecdotes from hands-on use: "It churned away for 17m51s and built me several .blend files." He also notes OpenAI’s claim that their image systems are used for "more than 3 billion images across ChatGPT Images and the GPT‑Image models in the API" and summarizes their guidance: "Choose Sunburst for workflows where editing precision matters most, and Flare for fast, high-quality everyday image generation."
Willison also flags a serious security demonstration from Calif Research: "Working with AI, our team found the bug and wrote the first remote code execution (RCE) exploit in about two days. Building the worm took one more week." And he relays a controversy around claims that an unreleased OpenAI model contributed to a Navier–Stokes breakthrough, quoting Tristan Buckmaster: "I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI."
Dwarkesh Patel shifts attention to what drives progress: "How much of the rapid progress in AI that we’ve seen over the last few years1 has come from data versus model improvements?" He describes the method: "We train combinations of these year-representative model recipes and data corpuses across different scales of training compute (up to 1e19 FLOPs)2." Since cross-entropy cannot be used here, "we evaluate these models on end capabilities as measured by the OLMES eval (which aggregates 10 different relatively easy benchmarks, mostly multiple choice QA)."
Lambert’s other note points to a structural shift in the ecosystem: "In 2026, open models are more competitive than ever, which has led to two interesting developments: Western model makers adopt open licenses, with both Google and Meta switching to Apache 2.0."
Marcus, meanwhile, sees two institutional checks emerging: "Senator Blumenthal just sent a great letter to OpenAI, reminiscent of the famous Watergate-associated phrase, “What did the President know, and when did he know it?”" and a lawsuit from Protect Democracy. He worries regulatory "opacity also becomes a potential tool for authoritarianism" and adds, "I am worried, for example, that the criteria may pay insufficient attention to the importance of monitorability, one of the few tools we have to keep an eye on what AI agents are doing."
What our tracker shows right now
- Total open AI roles: 4,604 across 177 employers.
- By role family:
- Modelling and engineering: 1,861 open, 986 opened and 280 closed in 30 days, across 142 employers
- Data: 844 open, 503 opened and 150 closed in 30 days, across 125 employers
- Infrastructure: 400 open, 164 opened and 50 closed in 30 days, across 88 employers
- Research: 233 open, 70 opened and 16 closed in 30 days, across 52 employers
- Product and design: 95 open, 48 opened and 10 closed in 30 days, across 45 employers
- Evaluation and safety: 41 open, 17 opened and 4 closed in 30 days, across 19 employers
- Top employers:
- Accenture: 943 open, 712 opened in 30 days
- Capital One: 186 open, 113 opened in 30 days
- OpenAI: 162 open, 83 opened in 30 days
- Amgen: 158 open, 40 opened in 30 days
- Anthropic: 127 open, 57 opened in 30 days
- PwC: 125 open, 92 opened in 30 days
- Waymo: 95 open, 26 opened in 30 days
- Databricks: 85 open, 35 opened in 30 days
Where the posts match the hiring pulse
-
Underinvestment in oversight. Mowshowitz’s concern that monitorability is getting harder and Marcus’s call for accountability line up with the thin staffing in evaluation and safety. Only 41 evaluation and safety roles are open across 19 employers, versus 1,861 in modelling and engineering. Our numbers do not show a broad pivot into safety hiring.
-
Benefits still feel back-end. Lambert’s claim that consumer touchpoints are small is consistent with the mix we see. Product and design stands at just 95 open roles across 45 employers. Meanwhile, integrators and services are ramping: Accenture alone lists 943 open roles, with 712 opened in 30 days. That hiring profile suggests enterprises are busy wiring models into workflows rather than shipping a wave of new consumer-facing products.
-
Data remains central. Patel’s thesis that progress is heavily data-driven finds support in demand for data talent: 844 open data roles (503 opened in 30 days) make it the second largest family behind modelling and engineering. Employers are staffing data pipelines and curation alongside model building.
-
Capabilities push continues. Raschka’s and Willison’s hands-on reports of stronger 3D and image workflows track with continued research and infrastructure investment: 233 research roles and 400 infrastructure roles are open.
Where our data do not support the claims (or cannot test them)
-
Boycott effects. Marcus’s "The time for a public boycott has come" is not yet visible in hiring. OpenAI lists 162 open roles with 83 opened in 30 days; Anthropic has 127 open with 57 opened. Those figures suggest continued expansion at the very labs under fire.
-
Mass consumer uptake. Willison relays usage claims of "more than 3 billion images across ChatGPT Images and the GPT‑Image models in the API." Our tracker cannot validate usage, but the relatively small product and design hiring pool suggests the jobs market has not swung toward consumer productization at the same pace as core engineering.
-
Security inflection. Calif Research’s "remote code execution (RCE) exploit in about two days" with AI is alarming. Our role taxonomy does not break out security-specific roles today, so we cannot say whether employers are staffing up directly in response to this class of risk.
-
Specific research controversies. Claims around an unreleased model’s role in a Navier–Stokes breakthrough are beyond the scope of hiring data. Our numbers neither confirm nor contest Willison’s account or the quotes he includes from Tristan Buckmaster.
What to watch next
-
Regulatory pressure and safety headcount. Marcus’s note about Senator Blumenthal’s letter and litigation over evaluation criteria raise the question Marcus himself asks: whether "the criteria may pay insufficient attention to the importance of monitorability." If that changes, we would expect the 41 open evaluation and safety roles to grow and diversify across more than 19 employers.
-
Enterprise-first demand. If Lambert is right that average users still see small benefits, the current skew to Accenture, PwC and similar employers is likely to persist. Watch whether product and design openings rise from today’s 95 as firms push polished AI features to end users.
-
Data hiring vs. model hiring. Patel’s data-centric view implies continued demand for data roles. Today’s 844 open data roles already outnumber research and infrastructure combined; if that gap widens, it would reinforce his conclusion about where progress is coming from.
-
The bigger arc. Our jobs-created indicator is counted from employer listings and calibrated to the World Economic Forum’s Future of Jobs Report 2025 projection of 11 million AI roles created and 9 million displaced by 2030. The present mix of roles suggests the near-term emphasis is on building and integrating systems, not on oversight. Whether that balance shifts will be the most important hiring signal to watch in the weeks ahead.
What we read
Every quote above is taken verbatim from one of these posts.
- Nathan Lambert: When will average people feel AI’s impact?, Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open mo
- Gary Marcus: The Case for Boycotting Generative AI, OpenAI’s Egregious Pattern of Misconduct, Two dire warnings, one from Terence Tao, the other from someone who ju, BREAKING: Two rays of hope
- Zvi Mowshowitz: Astra Is Hard to Monitor, GPT-6 Astra: The System Card, Alignment and What Comes Next
- Sebastian Raschka: GPT-6 Astra, Looped Transformers, and Hidden Reasoning
- Simon Willison: On the Navier–Stokes Millennium Prize Problem, Introducing ChatGPT Images 2.5, Quoting Calif Research, .blend URL Viewer
- Dwarkesh Patel: Pretraining progress is mostly coming from data