Frontier AI has outgrown the lab. The decisive questions now are about power — who builds the models, who controls them, and who gets to build on top of them. AI Frontiers is for the people doing the building: founders and operators creating products, companies, and strategy at the edge of what AI can do — on infrastructure owned by a handful of labs and governed from a handful of capitals. Each season charts where that frontier has moved, from the labs shipping the models to the capitals writing the rules, and what it means for anyone building something that lasts on ground that keeps shifting. Hosted by Fabio Lauria, founder of ELECTE. No hype, no jargon — strategy, stakes, and a builder's-eye view of the most consequential infrastructure of the century.
The Illusion of Reasoning: The Debate Shaking the World of AI
•Fabio Lauria•Episode 36
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
0:00
|
8:34
Unravel the enigma of artificial intelligence in our latest episode, where we dive into the heated debate ignited by Apple's groundbreaking research on the limitations of Large Language Models (LLMs). Are these AI systems truly capable of reasoning, or are we witnessing an illusion crafted by complex algorithms? Join us as we explore the implications of this debate on the future of AI, examining the boundaries of machine intelligence and the potential impact on industries and society. Key topics include the nature of reasoning in AI, the challenges faced by LLMs, and the broader implications for enterprise strategy and innovation. This episode promises to challenge your understanding of AI's capabilities and limitations. Tune in to gain insights into the evolving landscape of artificial intelligence and its role in shaping our digital future. Don't miss this thought-provoking discussion that could redefine your perspective on AI's potential.
In recent months, the artificial intelligence community has been embroiled in a heated debate sparked by two influential research papers published by Apple. The first, GSM Symbolic, October 2024, and the second, The Illusion of Thinking, June 2025, question the supposed reasoning abilities of large language models, sparking mixed reactions across the industry. As already analyzed in our previous in-depth article on the illusion of progress, simulating general artificial intelligence without achieving it, the question of artificial reasoning touches on the very heart of what we consider intelligence in machines. What Apple's researchist says Apple researchers conducted a systematic analysis of large reasoning models, LRMs, models that generate detailed reasoning traces before providing an answer. The results were surprising and for many alarming. The tests conducted, the study subjected the most advanced models to classic algorithmic puzzles such as Tower of Hanoi, a mathematical puzzle first solved in 1957 river crossing problems, logical puzzles with specific constraints, GSM symbolic, benchmark, variations of elementary level mathematical problems, controversial results. The results show that even small changes in the formulation of the problems lead to significant variations in performance, suggesting a worrying fragility in reasoning. As reported in Apple Insider's coverage, the performance of all models declines when only the numerical values in the GSM symbolic benchmark questions are altered. The counteroffensive, the illusion of the illusion of thinking. The AI community's response was swift. Alex Lawson of Open Philanthropy, in collaboration with Claude Opus of Anthropic, published a detailed rebuttal titled The Illusion of the Illusion of Thinking, challenging the methodologies and conclusions of the Apple study. The main objections ignored output limits. Many failures attributed to reasoning collapse were actually due to the model's output token limits. two incorrect evaluation. The automated scripts classified even partially correct outputs as total failures. Three, impossible problems. Some puzzles were mathematically unsolvable, but the models were penalized for not solving them. Confirmation tests. When Lawson repeated the test using alternative methodologies, asking the models to generate recursive functions instead of listing all moves, the results were dramatically different. Models such as Claude, Gemini, and GPT correctly solved 15 disc tower of Hanoi problems well beyond the complexity where Apple reported zero successes. Authoritative voices in the debate Gary Marcus. The longtime critic Gary Marcus, a longtime critic of LLM reasoning abilities, embraced Apple's findings as confirmation of his two decade old thesis. According to Marcus, LLMs continue to struggle with distribution shift, the ability to generalize beyond training data, remaining good solvers of problems that have already been solved. The local llama community, the discussion has also spread to specialized communities such as local llama on Reddit, where developers and researchers debate the practical implications for open source models and local implementation. Beyond the controversy, what it means for businesses, strategic implications, this debate is not purely academic. It has direct implications for AI deployment in production. How much can we trust models for critical tasks? RD investments. Where should we focus resources for the next breakthrough? Communication with stakeholders. How can we manage realistic expectations about AI capabilities? The neurosymbolic approach. As highlighted in several technical insights, there is an increasingly clear need for hybrid approaches that combine neural networks for pattern recognition and language understanding, symbolic systems for algorithmic reasoning and formal logic timing in strategic context, has not escaped observers that Apple's paper was published shortly before WWDC, raising questions about strategic motivations. As noted in the analysis by 9 to 5 Mac, the timing of Apple's paper, just before WWDC, raised some eyebrows. Was this a research milestone or a strategic move to reposition Apple in the broader AI landscape? Lessons for the future for researchers. Experimental design. The importance of distinguishing between architectural limitations and implementation constraints, rigorous evaluation, the need for sophisticated benchmarks that separate cognitive capabilities from practical constraints, methodological transparency, the obligation to fully document experimental setups and limitations for companies, realistic expectations, recognizing current limitations without giving up on future potential hybrid approaches, investing in solutions that combine the strengths of different technologies, continuous evaluation, implement testing systems that reflect real world usage scenarios, conclusions, navigating uncertainty. The debate sparked by Apple's papers reminds us that we are still in the early stages of understanding artificial intelligence. As highlighted in our previous article, the distinction between simulation and authentic reasoning remains one of the most complex challenges of our time. The real lesson is not whether LLMs can reason in the human sense of the word, but rather how we can build systems that leverage their strengths while compensating for their limitations. In a world where AI is already transforming entire industries, the question is no longer whether these tools are intelligent, but how to use them effectively and responsibly. The future of enterprise AI will likely lie not in a single revolutionary approach, but in the intelligent orchestration of several complementary technologies. And in this scenario, the ability to critically and honestly assess the capabilities of our tools becomes a competitive advantage in itself. For insights into your organization's AI strategy and the implementation of robust solutions, our team of experts is available for personalized consultations, sources and references, GSM symbolic, understanding the limitations of mathematical reasoning in large language models, Apple Machine Learning Research, The Illusion of Thinking, Understanding the Strengths and Limitations of Reasoning Models, Apple Machine Learning Research New Paper pushes back on Apple's LLM reasoning collapse study. Nike to five Max 7 replies to the viral Apple Reasoning Paper, Gary Marcus, The Illusion of Thinking. What the Apple AI Paper says about LLM reasoning. Arise AI. Apple's study proves that LLM based AI models are flawed. Apple Insider. The Illusion of Progress. Simulating General Artificial Intelligence Without Achieving It. Elect to share the newsletter. Welcome to Electe's newsletter, English. This newsletter explores the fascinating world of how companies are using AI to change the way they work. It shares interesting stories and discoveries about artificial intelligence and business, like how companies are using AI to make smarter decisions, what new AI tools are emerging, and how these changes affect our everyday lives. You don't need to be a tech expert to enjoy it. It's written for anyone curious about how AI is shaping the future of business and work, whether you're interested in learning about the latest AI breakthroughs, understanding how companies are becoming more innovative, or just want to stay informed about tech trends. This newsletter breaks it all down in an engaging, easy to understand way. It's like having a friendly guide who keeps you in the loop about the most interesting developments in business technology without getting too technical or complicated. Subscribe now, subscribe to get full access to the newsletter and publication archives.