Practitioners flag brittle agents and broad AI risks; our tracker shows employers mostly hiring builders and infra, with scant evaluation and safety roles.
Builders push forward while safety alarms get louder
Across the past two days, practitioners drew a sharp contrast: coding agents are both increasingly capable and worryingly fragile, and society needs a plan. Simon Willison warns that "Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attacks. They recently made that the default and have made bold claims about its effectiveness." His write-up highlights a successful attack found by Johann Rehberger: "He found an attack against auto mode which he claims works 80% of the time, by tricking Claude Code into downloading and uncompressing a zip archive, then executing code that imports base64 without noticing that this will import and execute a local struct.py file extracted from the archive." Willison documents that in response, "In a few runs Claude tried to terminate the malware process once it noticed the compromise, but Auto Mode denied the cleanup command." His bottom line is operational: "I agree with Johann's conclusion here: the only safe way to run agents if there's any risk of attracting the attention of an adversarial attack is with a sandbox: Run unattended coding agents in a container, VM or OS sandbox. Restrict network egress."
At the same time, Gary Marcus amplified employment and societal risk themes from Bill Gates, including that "Even criminals with very limited skills will be able to target victims at every scale" and the concern that "AI could stunt our kids’ development and replace human relationships" alongside a call for "creating a domestic and international framework for dealing with AI" and for "rebalanc[ing] how we tax labor and capital."
On capability and speed, Ethan Mollick contrasts the slow-adoption story with rapid industrial change: "I hear the story about how it took 30 years to gain productivity from electricity a lot in relation to AI (I have even told it myself) But not every technology is that way. Ford went from inventing the modern assembly line to rebuilding all work around it in 3 years, cutting car making time by 88%." And Simon Willison relays Paul Dix’s claim about large-scale AI-assisted software production: "The fact that AI wrote 1M LOC and then refined it over the course of the next couple of months to produce a reliable piece of software that is currently running on millions of developer machines is absolutely mind blowing. And you can say, “well it’s not that impressive because they had an oracle to compare against, so it was simple to go from one language to another”, but I think that’s selling this entire thing short. If you can build a verification system and give proper direction, AI can produce a highly complex, highly sophisticated piece of software and it can continue to refine it until it just works."
Zvi Mowshowitz captures the backdrop: "Cyber Lack of Security. Chinese hackers broke into the Federal Reserve?" and the employment anxiety: "They Took Our Jobs. Bill Gates warns of ‘economic catastrophe’ and more." He also urges readers to resist deferring blindly to status: "Think for yourself, schmuck. Or, as I once put it: You Have The Right To Think, also the moral duty to do so."
We tested these claims against our hiring tracker. Where they align, we say so. Where they do not, we say that plainly.
Hiring today is weighted to builders, not evaluators
Our tracker counts 4,294 open AI roles across 175 employers, taken from their own job feeds. The balance of those openings does not support the idea that evaluation and safety work is a near-term hiring priority, at least in job postings. By role family:
- Modelling and engineering: 1,651 open, 545 opened and 101 closed in 30 days, across 141 employers
- Data: 803 open, 274 opened and 41 closed in 30 days, across 124 employers
- Infrastructure: 403 open, 91 opened and 23 closed in 30 days, across 86 employers
- Research: 227 open, 43 opened and 7 closed in 30 days, across 51 employers
- Product and design: 101 open, 27 opened and 2 closed in 30 days, across 50 employers
- Evaluation and safety: 40 open, 12 opened and 3 closed in 30 days, across 19 employers
Relative to the concerns voiced by Marcus (via Gates) and the concrete agent failures described by Willison, just 40 evaluation and safety roles open across 19 employers is a small footprint. That is a clear mismatch: the commentary emphasizes urgent risks and governance, but current hiring skews to engineering and data.
Agent fragility vs infrastructure hiring
Willison’s recommendation to "Run unattended coding agents in a container, VM or OS sandbox. Restrict network egress" implies a practical pivot to hardened environments and operational controls. Our data shows 403 open infrastructure roles, with 91 opened and 23 closed in the past 30 days, across 86 employers. That is material demand and partially consistent with the call to put agents behind sandboxes and network restrictions. If organizations are standing up secure runtimes, observability, and isolation for agentic systems, they likely do that through infrastructure hiring.
But that same alignment does not extend to specialized safety and evaluation roles. With just 40 openings in that family, employers are not, at least in their job posts, moving at the pace of Willison’s and Marcus’s caution. Zvi’s note about "Cyber Lack of Security" fits the risk narrative, but job postings still concentrate on building and shipping.
Speed of adoption, reflected in openings
Mollick’s Ford analogy argues that adoption can compress from decades to a few years. In our tracker, the near-term cadence of openings does show momentum on the build side: 545 modelling and engineering roles opened in the last 30 days and 274 data roles opened. That is consistent with Paul Dix’s claim (as quoted by Willison) that with the right scaffolding "AI can produce a highly complex, highly sophisticated piece of software and it can continue to refine it until it just works." Employers appear to be hiring the people who will provide that scaffolding and integration.
Top employers by current openings underscore this point:
- Accenture: 638 open, 365 opened in 30 days
- Capital One: 188 open, 57 opened in 30 days
- Amgen: 175 open, 25 opened in 30 days
- OpenAI: 154 open, 66 opened in 30 days
- Anthropic: 121 open, 40 opened in 30 days
- PwC: 113 open, 48 opened in 30 days
- Waymo: 96 open, 15 opened in 30 days
- Databricks: 92 open, 22 opened in 30 days
Consulting, finance, biotech, foundation model labs, and autonomous systems are all posting. That breadth supports Mollick’s point that not every technology follows a slow S-curve.
Jobs, risk, and the near-term signal
Marcus’s summary of Gates includes the need for "creating a domestic and international framework for dealing with AI" and for "rebalanc[ing] how we tax labor and capital." Zvi flags Gates’s warning of "‘economic catastrophe’" in a section header. Our hiring figures cannot adjudicate long-run macro risk. They can say what employers are doing this month.
They are hiring. Our headline jobs-created indicator, counted from employer job listings and calibrated to the World Economic Forum’s Future of Jobs Report 2025 (11M AI roles created, 9M displaced, by 2030), is reflected in the 4,294 open roles we see today. That does not disprove displacement pressures, and it does not validate or invalidate any tax policy proposal. It does contradict an imminent freeze in AI hiring. The short-run signal is continued investment in building teams.
One more cautionary note from Mollick’s new work on agents: "🚨Our new research examines agentic shopping: can you consistently predict (or, using marketing, influence) what an agent chooses? Nope. We found that even small differences (viewing order of pages, memories) changed AI preferences in unpredictable ways. papers.ssrn.com/sol3/papers...." That unpredictability argues for more evaluation and safety capacity. Yet, as above, our postings show only 40 such roles open.
What this means for the next quarter
- Security and evaluation are under-represented in postings relative to the risks practitioners are documenting. Willison’s concrete agent failure case and Marcus’s risk list are not matched by a surge in evaluation and safety hiring.
- Infrastructure demand is healthy. That partially aligns with Willison’s operational advice to sandbox agents.
- Builder roles dominate. Mollick’s and Paul Dix’s capability and speed claims are consistent with 819 openings opened in the last 30 days across modelling and engineering plus data alone.
Zvi’s epistemic warning applies here: "Think for yourself, schmuck." The narrative is contested; the jobs data provide a ground truth about what employers prioritize now. If the risk conversation starts to drive budgets, we should see evaluation and safety postings rise. Until then, hiring remains concentrated on people who can ship, run, and integrate AI systems.
What we read
Every quote above is taken verbatim from one of these posts.
- Zvi Mowshowitz: AI #183: Pre Post Mortem, Against Modesty’s Bailey
- Simon Willison: Breaking Claude Code Opus 5 Auto Mode, Quoting Paul Dix, Qwen3.8-Flash-Next
- Noah Smith: The death of Market Street
- Ethan Mollick: I hear the story about how it took 30 years to gain productivity from , I know it is a lot for instructors who never asked for AI to change th, 🚨Our new research examines agentic shopping: can you consistently pre
- Gary Marcus: Excellent new Bill Gates essay on the urgency of having a coherent AI
- Emily M. Bender: Help! I was listening to a podcast episode recently where the host and
- David Ha: Hugging Face 🤝 NVIDIA