Mercor found AI faster and more accurate than 12 licensed CPAs on simplified bookkeeping tasks. On the full 160-task APEX benchmark, Claude Opus 5.5 led with 61.8%, while no model fully solved nearly 60% of tasks, leaving AI unable to close books without human oversight.