3 Sources
[1]
OpenAI Researchers Find That Even the Best AI Is "Unable To Solve the Majority" of Coding Problems
OpenAI researchers have admitted that even the most advanced AI models still are no match for human coders -- even though CEO Sam Altman insists they will be able to beat "low-level" software engineers by the end of this year. In a new paper, the company's researchers found that even frontier
[2]
AI can fix bugs -- but can't find them: OpenAI's study highlights limits of LLMs in software engineering
Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Large language models (LLMs) may have changed software development, but enterprises will need to think twice about entirely replacing human software engineers with LLMs,
[3]
OpenAI Thinks LLMs Can Earn $1M from Freelance Software Engineering Tasks
OpenAI has introduced SWELancer, a new benchmark to test whether frontier large language models (LLMs) can successfully complete real-world freelance software engineering tasks -- and even earn up to $1 million in total payouts. The evaluation is based on 1,488 freelance software engineering jobs
Share
Copy Link
OpenAI researchers develop a new benchmark called SWE-Lancer to test AI models' performance on real-world software engineering tasks, revealing that even advanced AI struggles with complex coding problems.

OpenAI researchers have developed a new benchmark called SWE-Lancer to evaluate the performance of large language models (LLMs) in real-world software engineering tasks. This innovative benchmark, based on over 1,400 freelance software engineering tasks from Upwork, aims to test the capabilities of frontier AI models in coding and software development
1
3
.SWE-Lancer comprises two main categories of tasks:
The benchmark includes projects ranging from quick $50 bug fixes to complex $32,000 feature implementations, with a cumulative value of approximately $1 million
3
. To ensure a fair assessment, the AI models were not allowed internet access during the tests, preventing them from simply copying existing solutions1
.Three advanced LLMs were put to the test using the SWE-Lancer benchmark:
The results revealed significant limitations in the AI models' abilities to handle complex software engineering tasks:
2
.2
.1
.Related Stories
The study's findings have important implications for the future of AI in software development:
1
.1
2
.2
.This research comes at a time when the impact of AI on various industries, including software development, is being closely scrutinized. A recent survey by Anthropic revealed that approximately 36% of all occupations incorporate AI for at least a quarter of their tasks, with software development being a key area of AI utilization
3
.As AI technology continues to advance, it's clear that while it can be a powerful tool for augmenting human capabilities in software engineering, it is not yet ready to fully replace human expertise. The SWE-Lancer benchmark provides a valuable tool for assessing progress in this field and understanding the economic implications of AI in software development
3
.Summarized by
Navi
[1]
[2]
1
Policy and Regulation

2
Technology

3
Technology
