Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
FlowerBench: Benchmarking AI Agents on Real Enterprise Work (flower.ai)
5 points by dimitrisflwr 21 days ago | hide | past | favorite | 1 comment


AI agents are advancing quickly, but enterprise work remains a much harder test for agents.

Today coinciding with ICML, we are introducing FlowerBench: a frontier agent benchmark for evaluating AI agents on secure, proprietary, long-horizon enterprise tasks.

FlowerBench evaluates AI agents on real enterprise work while always keeping private data, tools, and context inside the organization. Evaluations are coordinated across the Flower Enterprise Evaluation Network, where only sanitized non-sensitive results are shared.

We have already assessed proprietary and open agents and models across a range of enterprise-grade tasks in various industries: Healthcare, Insurance, Operations, MLOps, Legal, Marketing and Finance.

Looking forward to feedback from people working on agent benchmarks, enterprise AI, tool-use evals, and long-horizon tasks evaluation.

*Other Useful Links*

Blogpost: https://flower.ai/blog/2026-07-06-flowerbench

Share your Agent Task: https://flowerlabs.typeform.com/to/AKGDGUhf




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: