9 Sources
[1]
OpenAI says GPT-5 stacks up to humans in a wide range of jobs | TechCrunch
OpenAI released a new benchmark on Thursday that tests how its AI models perform compared to human professionals across a wide range of industries and jobs. The test, GDPval, is an early attempt at understanding how close OpenAI's systems are to outperforming humans at economically valuable work --
[2]
OpenAI Says ChatGPT Can Already Do Some Work Tasks as Well as Humans
OpenAI is trying to make the case that AI can actually be useful at work, as some recent studies have shown that companies aren’t getting much out of their AI investments. On Tuesday, the ChatGPT-maker released a report introducing a new benchmark for testing AI on “economically valuable,
[3]
OpenAI is now testing ChatGPT against humans in 44 different occupations, from lawyers and software developers to registered nurses -- here's the full list of jobs affected
OpenAI, the company behind ChatGPT, has announced a new benchmark for testing its GPT-5 model, which involves pitting the AI directly against human experts in a variety of occupations. The benchmark is called GDPval and is responsible for assessing how close ChatGPT is getting to outperforming
[4]
OpenAI tool shows AI catching up to human work
Why it matters: We're at an AI reckoning, where leaders are trying to justify investments without effective tools to measure returns. * A recent MIT study showing that most AI projects fail launched a debate about its techniques, but also exposed the challenges in measuring returns on these
[5]
OpenAI Releases List of Work Tasks It Says ChatGPT Can Already Replace
"Today's best frontier models are already approaching the quality of work produced by industry experts." ChatGPT maker OpenAI has released a new evaluation, dubbed GDPval, to measure how well its AIs perform on "economically valuable, real-world tasks across 44 occupations." "People often
[6]
AI Isn't Taking Your Job Yet -- But It Might Soon, OpenAI Data Suggests
The study showed the first wave of disruption will hit office-based jobs, from coders to lawyers and journalists. OpenAI unveiled GDPval on Thursday -- a benchmark that tries to assess qualitatively whether AI can do your actual job. These are not hypothetical exam questions, but real
[7]
OpenAI: GDPval framework tests AI on real-world jobs
OpenAI has announced a new evaluation framework, GDPval, to measure artificial intelligence performance on economically valuable tasks. The system tests models on 1,320 real-world job assignments to bridge the gap between academic benchmarks and practical application. The GDPval framework
[8]
OpenAI Tests if ChatGPT 5 Can Automate Your Job With Unexpected Findings
What if the future of your job wasn't about being replaced by AI, but about working alongside it? The rapid advancements of tools like GPT-5 have sparked both excitement and anxiety, with many wondering whether machines will soon outperform humans in the workplace. OpenAI's latest research,
[9]
OpenAI's GPT-5 matches human performance in jobs: What it means for work and AI
On September 25, 2025, OpenAI dropped a bombshell: its latest model, GPT-5, now "stacks up to humans in a wide range of jobs." The declaration ripples far beyond the world of AI benchmarks, it raises urgent questions about the future of work, the boundary between human and machine, and how
Share
Copy Link
OpenAI introduces GDPval, a new benchmark to evaluate AI performance across 44 occupations. Results show top AI models, including GPT-5 and Claude Opus 4.1, are nearing human expert-level quality in many tasks.
OpenAI has unveiled a new benchmark called GDPval, designed to evaluate the performance of AI models on 'economically valuable, real-world tasks' across 44 different occupations
1
2
. This benchmark aims to ground conversations about AI's impact on the workforce in evidence rather than speculation, and to track model improvements over time4
.
Source: Digit
GDPval focuses on nine industries that contribute significantly to the U.S. Gross Domestic Product (GDP)
1
4
. The benchmark includes around 1,300 specialized tasks crafted by experienced professionals with an average of 14 years of experience3
. These tasks span various deliverables such as legal briefs, engineering blueprints, customer support conversations, and nursing care plans2
3
.
Source: Decrypt
OpenAI's tests revealed that leading AI models are approaching parity with human professionals on many tasks
4
. Notably:4
.4
.1
.
Source: Axios
The AI models showed varying levels of proficiency across different jobs:
2
.2
.While the results are promising, OpenAI emphasizes that AI is not poised to replace humans entirely
2
5
. Instead, the company suggests that AI could complement human workers, allowing them to focus on more creative and judgment-intensive aspects of their jobs1
4
.Related Stories
OpenAI acknowledges that GDPval is an early step and doesn't capture the full complexity of many economic tasks
3
. Future versions of the benchmark are expected to include more interactive workflows and context-rich tasks to better reflect real-world knowledge work3
.The introduction of GDPval comes at a time when the AI industry is facing scrutiny over the practical value of AI investments. A recent MIT study found that fewer than one in ten AI pilot projects delivered measurable revenue gains
2
. Critics have also raised concerns about 'workslop' – AI-generated content that appears good but lacks substance2
.As AI continues to evolve, its impact on the job market remains a topic of intense debate. While OpenAI's research suggests significant progress in AI capabilities, the full implications for various industries and occupations are yet to be fully understood.
Summarized by
Navi
15 Sept 2025•Technology

09 Jul 2026•Technology

02 Dec 2025•Technology

1
Technology

2
Technology

3
Policy and Regulation
