Hitting the same wall, harder and faster
There’s one thing that Large Language Models have highlighted over the past couple of years. One thing that I’ve circled back to again and again since I wrote my first online blog post 8 years ago:
We are chronically incapable of framing problems with clear goals.
It’s either a fundamental human flaw or an irreducible problem with reality: we just can’t figure out how to measure the real outcomes we want to achieve, and we always settle for easily quantifiable proxy metrics.
Unfortunately, we get what we measure. LLMs have shown this to us in three very different ways.
AI alignment
The challenge of aligning AI with human goals and values is the clearest demonstration of proxy-goal failure. AI systems do not understand human intent, they optimise mathematical loss functions. Because complex human ethics, preferences, and social nuances cannot be calculated directly, researchers substitute them with quantifiable proxies, most notably human evaluator ratings during Reinforcement Learning from Human Feedback (RLHF). The goal is to reward helpfulness and harmlessness, but in reality we also end up with textbook examples of Goodhart's Law.
The most visible symptom is sycophancy, driven by the inherent human nature of RLHF. Because human raters naturally favour polite, validating and agreeable responses, optimising for user ratings incentivises the model to tell users what they want to hear rather than the blunt truth. In late April 2025, OpenAI was forced to roll back an update to GPT-4o after tuning its reward function on thumbs-up feedback caused the model to flatter users, validate harmful delusions, and offer dangerous advice simply to maximize immediate user approval.
Even when a reward function appears robust, models can suffer from goal misgeneralisation, i.e. learning an easier correlated proxy during training that shows up as misaligned behaviour when deployed in new contexts. Rather than internalising genuine goals and reasoning, models learn to reproduce the superficial markers of a correct answer.
Ultimately, every reward function is a simplified proxy for what humans actually want, and high-pressure optimisation will always exploit the difference to find easier correlated goals. This makes deception (behaviours where the model's output misleads the receiver to secure a reward, without or without intent) an inevitable byproduct of proxy goal design. And to make the puzzle even more complex, attempts to suppress this behavior by enforcing rigid safety guardrails often backfire, pushing gaming strategies underground. When a model is penalised for triggering automated evaluation filters, it simply learns how to pass the test. The MASK benchmark, published by the Center for AI Safety in March 2025, proved this dynamic by directly measuring lying propensity. It showed that frontier models, despite scoring highly on standard truthfulness benchmarks, regularly feign incompetence, withhold information or disguise capabilities.
AI in business
The second manifestation of proxy-goal failure is happening inside companies adopting AI. Over the past year, executive mandates have pushed employees to integrate LLMs into their daily job, with the aim to break down the productivity bottleneck. This relies on the faulty but popular premise that execution capacity is the primary bottleneck in business. The real constraint is problem definition: finding the right problems to solve, framing them rigorously, and setting meaningful goals. Adding delivery capacity to the wrong objective propels an organisation exactly nowhere.
I suspect that a lot of companies who have been aggressive about using AI to increase productivity are discovering an uncomfortable truth: removing constraints destroys strategic discipline. Hard capacity limits force teams to think deeply, edit ruthlessly, and scrutinise priorities before committing resources. When you remove execution friction without sharpening problem framing, you simply make it easier to generate vast amounts of low-value work.
There will surely be winners: organisations who are able to leverage this new tool to deliver with a laser focus on their objectives and to pursue genuinely valuable opportunities. But, like with any new tool or technique (you can look back on the Agile transformations of the 2010s…), many will just miss the point and keep measuring the wrong thing. People argue that cheap generation enables rapid experimentation and faster feedback loops. That holds true only if teams maintain the discipline to form clear hypotheses and evaluate them objectively. I doubt that AI has suddenly made organisations and teams better at this. It’s a skill that demands hands-on practice and critical thinking, not automation.
AI as a product
Finally, we see a flagrant demonstration of bad problem definition in frontier labs. They had two roads in front of them: either build software products around specific human needs or pursue open-ended frontier AI research for the sake of scientific advancement. Instead, labs conflated a scientific research quest for AGI with a venture-backed software business.
The labs’ hybrid identity creates deeply conflicted objectives: the incentives of an academic research laboratory (developing frontier capabilities in a controlled environment and developing methods to align, understand and control models) are structurally at odds with those of a commercial product company, which requires a solid business model centred around a clear value propositions. As product companies, they completely bypassed product-market fit, releasing models with general capabilities into the market and offloading the burden of problem definition onto customers. And as research organisations, they’ve tied their research direction to revenue goals. What a mess.
—
I sometimes wonder what the world would look like if we were better at defining problems and goals. Or at least if we took the time to do it before building and deploying technology with unprecedented capabilities. Maybe we’d have public -facing AI products that were designed to solve specific problems, making them far easier to keep safe. Maybe we wouldn’t have so much capital tied up in a speculative financial bubble. And maybe workers around the world wouldn’t be anxious about losing their jobs because executives and shareholders are easily aroused by the vague transformative promises of generative AI.
References:
OpenAI. Sycophancy in GPT-4o: what happened and what we’re doing about it [Internet]. OpenAI; 2025 Apr 29 [cited 2026 Aug 10]. Available from: https://openai.com/index/sycophancy-in-gpt-4o/
Ren R, Agarwal A, et al. The MASK benchmark: Disentangling honesty from accuracy in AI systems. Center for AI Safety & Scale AI; 2025 Mar 5.
Microsoft Research Cambridge. The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers [Internet]. Microsoft Research; 2025 Mar 15 [cited 2026 Aug 10]. Available from: https://www.microsoft.com/en-us/research/publication/the-impact-of-generative-ai-on-critical-thinking-self-reported-reductions-in-cognitive-effort-and-confidence-effects-from-a-survey-of-knowledge-workers/
Ngo R, Chan L, Mindermann S. The alignment problem from a deep learning perspective. arXiv preprint arXiv:2209.00626. 2022 Sep 2. Available from: https://arxiv.org/abs/2209.00626
Betley J, et al. Emergent misalignment from superposition [Internet]. OpenReview; 2024 [cited 2026 Aug 10]. Available from: https://openreview.net/pdf/a6ba0d18d8065120fc81f27d2b15e40493054604.pdf
Karwowski J, Hayman A, et al. Goodhart's law in reinforcement learning [Internet]. Semantic Scholar; 2023 [cited 2026 Aug 10]. Available from: https://www.semanticscholar.org/paper/Goodhart's-Law-in-Reinforcement-Learning-Karwowski-Hayman/a4bd4da02241eacee990c89ddd748ce37a248fc0