LLMs perform much worse on corporate Q&A tasks at realistic scale (thousands of documents) than on smaller benchmarks, and CorporateBench provides a standardized way to measure this performance gap without exposing real company data.
CorporateBench is a large-scale Q&A benchmark for testing how well language models handle real-world corporate document collections. It includes 230,000+ human-validated documents organized into synthetic companies of varying sizes, with questions requiring both information extraction and knowledge base querying.