61Signal
Score
F
FastCompanyby María José Gutiérrez ChávezSeptember 28, 2026

Even AI can’t perfectly assemble Ikea furniture—yet

The article highlights the challenges of assembling Ikea furniture and how these challenges serve as a benchmark for assessing AI intelligence. For brand strategy, this underscores the importance of understanding consumer pain points and leveraging AI to enhance user experiences, potentially transforming how brands like Ikea can support customers in complex assembly tasks.

◎ EmergingdigitalstrategyIkeaOpenaiAnthropic

FastCompany: Assembling Ikea furniture is not for the faint of heart. So much so that folklore suggests assembling a dresser, a bed, or anything from the Swedish brand can create enough tension it may even destroy a relationship . But beyond the stuff’s being a pain to build, it turns out that the problem-solving and spatial reasoning involved in building Malm dressers or Billy bookshelves are remarkably good benchmarks for assessing just how intelligent AI models are.

Aiden Ament and Greg Burnham, researchers at Epoch AI, an artificial intelligence research nonprofit, recently released a report detailing the Furniture Assembly Benchmark , a metric developed to test the actual intelligence of AI models. It centers on how well the model can identify errors in partially assembled Ikea furniture. At a time where tech companies love bragging about how intelligent their models are, it can be complicated to assess how “smart” each model actually is—particularly compared with others. Asking Chat GPT or Claude to spot errors in Ikea furniture assembly puts the models to the test in real-world applications.

And while no model got it right 100% of the time, it does seem they are quickly improving. The Ikea test The researchers first set out to develop a benchmark while on a company retreat, figuring out how to set up an experiment to test AI’s capabilities. “First, we tried acting like ‘the AI’s robot’ where we asked it what to do and tried to follow its directions somewhat literally. The results were often messy,” Burnham tells Fast Company . “I like the final idea better, which was due to my colleague Aiden Ament: asking it to spot mistakes,” Burnham says.

“This reflects the reality that it’s OK to make mistakes midassembly so long as you can catch and correct them.” Researchers used three Ikea products for the experiment, picking them based on varying degrees of difficulty according to Ikea’s own Complexity Index : the Ställ shoe rack for an easy level, Tonstad bed frame for a medium level, and the Gullaberg dresser as the most complex. They went about assembling each piece of furniture, intentionally making mistakes along the way, and taking 60 pictures of various steps to feed them to the various AI models from OpenAI, Anthropic, Google, Moonshot, and Alibaba.

In addition to the images, the researchers provided each AI model with a PDF of the assembly instructions, a tool to zoom in on each image, and a Python interpreter. Each AI model was tasked with identifying what mistakes were made by the user and at what point during assembly, and was scored accordingly against the benchmark. Additionally, the researchers asked the model to provide users with an explanation of each mistake identified. The benchmark isn’t just about assembling Ikea furniture. It’s intended to go beyond that, to measure how well these models could guide and troubleshoot more complex builds.

“If models can reason quickly and accurately enough in this domain, it is plausible that AI could soon provide reliable, real-time guidance for complex physical assembly and repair tasks,” Ament and Burnham wrote in their report. “We think this challenge serves as a good proxy for a number of economically important tasks that require similar visual and spatial reasoning, like fixing a car or repairing household appliances,” they wrote. Initially, the best scoring model on the Ikea test was Anthropic’s Claude Opus 4.5, scoring at just 28% accuracy. But within 10 months, OpenAI quickly caught up, with its GPT-6 Astra model reaching 80%.

Article truncated for readability. Read the full piece →

Intelligence PanelSignal score: 60.5 / 100
Primary Signal
Emerging
Building momentum — trajectory being tracked
Brand Impact
Medium
Impact score: 60/100 — moderate relevance to positioning decisions
Novelty
Moderate
Novelty: 50/100 — iterative development of an existing theme
Action Priority
Soon
Flag for the next strategic review cycle
Scoring Rationale

The article discusses the intersection of AI and consumer experience in a well-known context, making it relevant for brand strategy, but the concept of using AI to address consumer pain points is not entirely new.

60
Impact
weight 35%
50
Novelty
weight 30%
70
Relevance
weight 35%
Brands Mentioned
IIkeaOOpenaiAAnthropicGGoogleMMoonshotAAlibaba
Related SignalsAll Signals →