Broad AI benchmarks are good at telling us which models look strong on public leaderboards. They are less good at catching the quiet failures that break production systems. Grab Bench is our internal evaluation framework for testing AI models and agents on Grab-shaped work, with task-specific scoring, hidden cases, and row-level failure analysis so teams can see what actually went wrong.