Mercor’s cover photo
Mercor

Mercor

Software Development

San Francisco, California 839,572 followers

Organizing human intelligence to power the AI economy.

About us

We find the best experts in every professional domain and put their knowledge to work training frontier models. Through APEX, we measure whether those models can actually perform economically valuable work. We're also bringing that expertise to enterprises: deploying custom AI agents, staffing teams with vetted domain experts, and helping organizations encode their own knowledge into AI systems.

Website
mercor.com
Industry
Software Development
Company size
201-500 employees
Headquarters
San Francisco, California
Type
Privately Held
Founded
2023

Locations

Employees at Mercor

Updates

  • View organization page for Mercor

    839,572 followers

    APEX turns 1 today. Last year, we launched the AI Productivity Index (APEX) to answer whether frontier AI models can do professional work. Since then, we’ve created new benchmarks while model capabilities progress more rapidly than even the boldest predictions. APEX remains the industry standard for evaluating frontier AI on economically valuable work. Here’s how APEX has evolved over the past year. 𝗢𝗰𝘁 𝟮𝟬𝟮𝟱: 𝗔𝗣𝗘𝗫 Mercor introduces APEX, our first AI benchmark for testing investment banking, law, consulting, and medicine. 𝗝𝗮𝗻 𝟮𝟬𝟮𝟲: 𝗔𝗣𝗘𝗫-𝗔𝗴𝗲𝗻𝘁𝘀 Built with partners Box and Harvey, APEX-Agents evaluates AI agents on long-horizon tasks in investment banking, consulting, and corporate law. 𝗠𝗮𝗿 𝟮𝟬𝟮𝟲: 𝗔𝗣𝗘𝗫-𝗦𝗪𝗘 Co-developed with Cognition, APEX-SWE assesses real production engineering across integration and observability. 𝗝𝘂𝗹 𝟮𝟬𝟮𝟲: 𝗔𝗣𝗘𝗫-𝗔𝗰𝗰𝗼𝘂𝗻𝘁𝗶𝗻𝗴 Built with Ramp, APEX-Accounting measures whether models can close the books. 𝗦𝗲𝗽 𝟮𝟬𝟮𝟲: 𝗔𝗣𝗘𝗫-𝗔𝗴𝗲𝗻𝘁𝘀 𝟭.𝟭 Our first major update to APEX-Agents includes newly audited tasks, an improved judge, and enhancements to ensure leaderboard accuracy. Thank you to all our partners who built these benchmarks with us. And, to the Mercor experts, professional bankers, lawyers, consultants, doctors, engineers, and accountants, whose judgement and expertise ensure these benchmarks are realistic. The rate of AI progress is only increasing. Mercor is committed to maintaining and extending our APEX family of benchmarks as the industry standard for informing decisions about the AI frontier and its ability to do professional work.

  • View organization page for Mercor

    839,572 followers

    How do frontier AI models compare to accountants on real accounting work? We hired 12 licensed accountants to complete four month-end close tasks adapted from APEX-Accounting. They had to dig through a company's working files, find the right numbers, do the math, and deliver a table of results. We ran frontier models on the same tasks, with the same files and grader. On these tasks, models were faster and more accurate than every accountant in our study. 18 months ago, the best models fell short of the average accountant's ~37% score. Today, models ace the same tasks. Models are also more than an order of magnitude cheaper per rubric item than the humans. On the items humans tend to get right, the AI is about 34x faster. These results don't mean accountants are replaceable. The study specifically evaluated what AI is particularly good at like searching across files and following instructions closely. Real accounting work also involves understanding business-specific context and communicating with colleagues and clients. But the results do suggest substantial changes, and productivity gains, in accounting in the coming years. That would be true even if model progress were to stall today. Read the blog + full paper: https://lnkd.in/ghn_RzCY

    • No alternative text description for this image
  • View organization page for Mercor

    839,572 followers

    Mercor runs more than 2.5 million model eval workloads every week across 200+ models. At peak load, some trajectories were spending ~80% of their time waiting for compute. Our engineering team built a dynamic compute scheduler to get more out of the capacity we already had. It continuously reallocates concurrency across models based on demand, latency, and available capacity. Across our six busiest model lanes, throughput increased 6x to 20x without adding provider capacity. GPT-5.5 throughput increased 20x. Chris Setian & Tong Pan explain how we built it, including the scheduling problem, the architecture, and the results: https://lnkd.in/gtqzW3eg We’re hiring engineers to work on problems like this. See our open roles at the link in the comments. 

    • No alternative text description for this image
  • View organization page for Mercor

    839,572 followers

    Gemini 4 Argon is the new #1 on APEX-Agents. 82.2% Pass@1 (#1) 87.9% Mean score (#1) It is the first model to exceed 80% Pass@1. It is +6.7 pts over the prior #1, Sonnet 5.5, and +14.4 pts over the best prior model from DeepMind, Gemini 3.7 Flash (67.8%). #𝟭 𝗶𝗻 𝗮𝗹𝗹 𝗔𝗣𝗘𝗫-𝗔𝗴𝗲𝗻𝘁𝘀 𝗱𝗼𝗺𝗮𝗶𝗻𝘀 Google says Argon delivers frontier performance in "enterprise knowledge work like legal and finance." On APEX-Agents, Gemini 4 Argon ranks first in every professional domain. Management consulting: 90.3% (#1) Investment banking: 80.9% (#1) Corporate law: 75.3% (#1) Consulting appears to be Argon's strongest domain, scoring 10.3 pts ahead of Opus 5.5. It is also the first model to score above 90% on any APEX-Agents leaderboard. 𝗧𝗼𝗸𝗲𝗻 𝘂𝘀𝗮𝗴𝗲 𝗯𝘆 𝗔𝗣𝗘𝗫-𝗔𝗴𝗲𝗻𝘁𝘀 𝗱𝗼𝗺𝗮𝗶𝗻 Argon used 2.6M tokens per attempt on average. The median attempt used 1.75M. Investment banking: 3.5M per attempt Corporate law: 2.7M Management consulting: 1.7M Almost all Gemini 4 Argon’s token usage comes from input. It used 2.61M input tokens per attempt compared with only 28.5K output tokens per attempt. That is about 90 input tokens for every output token. The agent reads a lot of files before composing a short answer. Congrats to Google and Google DeepMind.

    • No alternative text description for this image
  • View organization page for Mercor

    839,572 followers

    GPT-6.1 Sol scores 60.0% Pass@1 and 73.2% mean on APEX-Agents (#11), and 46.9% on APEX-SWE (#18). OpenAI says it nearly matches GPT-6 Astra on professional work "at one-fifth of Astra's standard input and output token prices." That is consistent with what we’re seeing on APEX, within 4.7 points of Astra on Agents and 3.1 points on SWE. 𝗔𝗣𝗘𝗫-𝗔𝗴𝗲𝗻𝘁𝘀 GPT-6.1 Sol improves Pass@1 across all APEX-Agents domains compared with GPT-6 Sol. Corporate law: 63.4% (#11), up 7.2 Management consulting: 58.4% (#12), up 6.8 Investment banking: 58.1% (#10), up 2.9 That makes GPT-6.1 Sol OpenAI's #2 model on APEX-Agents, behind Astra and ahead of GPT-5.6 Terra (58.2%). It improves 5.7 points over GPT-6 Sol overall. On investment banking, GPT-6.1 Sol beats Astra (55.9%). It is now OpenAI's top model there. For corporate law, GPT-6.1 Sol trails Astra by 10 points on Pass@1, but its mean score is 85.3%. Only 4% of law runs score zero. It does most of the work and misses a rubric item or two. 𝗧𝗼𝗸𝗲𝗻𝘀 𝗮𝗻𝗱 𝗰𝗼𝗻𝘀𝗶𝘀𝘁𝗲𝗻𝗰𝘆 0.50M tokens per attempt on APEX-Agents. 1.29M per attempt on APEX-SWE. On law, failing runs use 2.4x the tokens of passing runs. GPT-6.1 Sol passed 49% of APEX-Agents tasks on all 4 runs. On APEX-SWE, 39%. 70 Agents tasks were never solved in any run. 𝗔𝗣𝗘𝗫-𝗦𝗪𝗘 GPT-6.1 Sol improves slightly over the prior generation with 46.9% (#18), up 1.9 from GPT-6 Sol. The gain is all in Observability where it scored 32.3% (#19), increasing by 4.5. Integration was mostly flat at 61.5%, scoring slightly below GPT-6 Sol (62.3%) and GPT-6 Astra (62.0%). Congrats OpenAI on the launch. See full leaderboards: https://lnkd.in/guwAacJQ

    • No alternative text description for this image
  • View organization page for Mercor

    839,572 followers

    Claude Sonnet 5.5 debuts at #1 on APEX-Agents and #2 on APEX-SWE. APEX-Agents: 75.5% Pass@1 (#1) 21 pt gain from from 54.5% for Sonnet 5 APEX-SWE: 66.4% Pass@1 (#2) 20 pt gain from 46.4% for Sonnet 5 Anthropic states Sonnet 5.5 at Max effort performs comparably to Opus 5.5. On APEX-Agents, it scores 2.0 pts higher (75.5% vs. 73.5%). APEX-Agents shows big improvement over Sonnet 5, where two domains gain more than 20 points. Investment banking: 77.2% Pass@1 (#1), up from 54.1% (+23.1 pts) Management consulting: 79.7% (#2), up from 47.8% (+31.9 pts) Corporate law: 69.6% (#5), up from 61.6% (+8.0 pts) In investment banking, Sonnet 5.5 takes over #1 from Gemini 3.7 Flash (71.3%). Consulting is the biggest jump over Sonnet 5, gaining +31.9 pts over the prior generation. On APEX-SWE, Sonnet 5.5 is 1.2 pts behind Opus 5.5 (67.6%). Integration: 69.3% (#1), up from 60.3% (+9.0 pts) Observability: 63.5% (tied #2), up from 32.5% (+31.0 pts) Sonnet 5 solved about 1 in 3 production debugging tasks, and now Sonnet 5.5 solves nearly 2 in 3. Congrats to Anthropic on the launch.

    • No alternative text description for this image
  • View organization page for Mercor

    839,572 followers

    Claude Opus 5.5 is the new #1 on APEX-SWE. Overall, Opus 5.5 scores 67.6% Pass@1, a +3.9 point gain over the previous leader, Opus 5, at 63.7%. APEX-SWE measures observability, debugging from production telemetry, and integration, if a model can build an end-to-end system. Compared with Opus 5: Observability Opus 5.5: 69.8% (#1) Opus 5: 63.5% (#2) It gains +6.3 points. Opus 5.5 leads the board by 6.3 over Opus 5 and 10.8 over Fable 5.1. Integration Opus 5.5: 65.5% (#4) Opus 5: 64.0% (#9) Opus 5.5 improved at reading telemetry and finding faults, but building systems from scratch moved less. Congrats to the Anthropic team.

    • No alternative text description for this image
  • View organization page for Mercor

    839,572 followers

    Mercor is committing $5M to a new AI Capabilities Fund. In conversations with researchers and enterprises, one challenge that comes up repeatedly is that many AI evals do not reflect how models are used. But to provide useful signal for eval and training, benchmarks need to capture actual workflows in representative environments. This will only become more difficult as models are used for more high-value, complex tasks. The Capabilities Fund supports researchers working on frontier eval challenges across a range of domains and data shapes, including: - Reward calibration and aligning grades with expert preferences - Creating high-quality realistic environments - Long-horizon workflows - Ambiguous requests and under-specified outcomes - Navigating social, temporal, and business context We will fund researcher time, API credits, and travel. We will also cover the cost to work with our network of 5M+ experts, as well as free use of our evals and analysis platform. Grants are available to independent researchers, non-profits, and academics. To submit an Expression of Interest, go to: https://bit.ly/3VeYFuw

    • No alternative text description for this image
  • View organization page for Mercor

    839,572 followers

    OpenAI GPT‑6 Sol and Luna are on the APEX leaderboards. GPT-6 Sol scores 54.3% Pass@1 (#18) on APEX-Agents. That is +2.9 pts over GPT-5.6 Sol. GPT-6 Luna scores 44.3% Pass@1 (#29), +1.3 pts over GPT-5.6 Luna. On APEX-SWE, Luna scores 38.8% Pass@1, +6.7 pts. OpenAI reduced API prices for Sol and Luna by 50% compared with their GPT‑5.6 promotional pricing. GPT-6 Sol improves in finance and consulting domains: Investment banking: 55.2% Pass@1 (#13), +5.5 pts over GPT-5.6 Sol Management consulting: 51.6% (#16), +9.0 pts GPT-6 Sol is only 0.7 pts behind GPT-6 Astra in investment banking. These gains are not consistent across all domains. Compared with GPT-5.6, Sol’s corporate law score drops 5.7 pts to 56.2% (#27). Luna drops 2.8 pts to 55.6%. GPT-6 Luna improves most on coding work: On APEX-SWE, GPT-6 Luna scores 38.8% Pass@1, gaining +6.7 pts over GPT-5.6. Integration: 59.3%, +13.3 pts Observability: 18.3%, no change Sol scores 45.0% Pass@1 (#18), about level with GPT-5.6 Sol (45.8%). Integration: 62.3% (#13), level with GPT-6 Astra (62.0%) Observability: 27.8% (#19), down 3.7 pts For both models, debugging with production telemetry is still the hardest part. Congrats to the OpenAI team on the launch.

    • No alternative text description for this image

Similar pages

Browse jobs