Benchmarking AI Agents on Real Workflows
Listen to the summary
Uses a voice available on your device
Audio options
On this page 4 sections
Related concepts 3 concepts
Key Takeaways
- Even the strongest tested model completes only 30% of StartupBench tasks successfully.
- The evaluation framework uses an automated Agent-as-a-Judge system that matches human expert judgment 92.78% of the time.
- Finance-related tasks are the most challenging, resulting in an average score of 54.48%.
- Current agents suffer from self-verification hallucination, where they incorrectly assume tasks are complete based on flawed internal reasoning.
Summary & Methodology Analysis
StartupBench addresses the gap between researcher-defined tasks and the requirements of real-world AI applications. The methodology involves surveying market-validated startups to define workflows, followed by expert reconstruction into natural-language requests with specific workspaces and weighted rubrics. The evaluation pipeline utilizes an automated Agent-as-a-Judge system to assess deliverables, ensuring high consistency with human experts at a 92.78% agreement rate. This approach shifts the focus from synthetic benchmarks to complex, deliverable-oriented tasks that mirror actual user expectations. The research included tests on models like GPT-5.6-sol, Kimi-K3, and DeepSeek-V4-Pro.
Interactive System Flowchart
Cross-Examination & FAQs
A deeper dive clarifying mechanics, constraints, and baseline evaluations.
Q1. What is the core purpose of StartupBench?
It serves as a benchmark for AI agents based on real-world, market-validated workflows rather than researcher-defined tasks.
Q2. How do agents generally perform on these tasks?
Performance is limited, with the strongest model completing only about 30% of the tasks successfully.
Q3. Are some industries harder for agents than others?
Yes, Finance is the most challenging domain, achieving an average score of 54.48%.
Q4. What is the 'Agent-as-a-Judge' evaluation method?
It is an automated, rubric-based evaluation system that assesses task deliverables across multiple dimensions independently.
Q5. How reliable is the automated scoring compared to human experts?
The automatic evaluator achieves 92.78% agreement with human expert judgments at the rubric level.
Q6. What is 'self-verification hallucination'?
It is a failure mode where the model assumes its reasoning process is correct and implies task success without verifying the actual output artifact.
Q7. How do models handle output formatting requirements?
No evaluated model achieves perfect compliance on the 56 tasks that have explicit output format requirements.
Q8. Do models succeed by focusing on core task requirements?
No, models frequently fail on core requirements even when they satisfy auxiliary ones, leading to deliverables that are practically unusable despite appearing complete.
Q9. What specific models were evaluated in this study?
Models included GPT-5.6-sol, GPT-5.5, Gemini-3.1-Pro, Seed-2.1-Pro, Qwen-3.6-Max, Kimi-K3, Kimi-K2.6, DeepSeek-V4-Pro, and GLM-5.1.