Collinear AI’s cover photo
Collinear AI

Collinear AI

Software Development

Shipping safer, smarter, and more reliable AI into enterprise environments.

About us

Collinear AI builds simulation labs where AI agents learn enterprise work before going to production. AI agents are inevitable. But they can't be trained or tested on the real world. You can't fix what you can't reproduce. They need a world to practice in -- that's what we build. Our labs give agents enterprise APIs, realistic users who push back and change their minds, tasks with real-world ambiguity, and verifiers that score what actually matters. So you can stress-test your agents before production and continuously improve them after. We work with frontier AI labs, and F500 enterprise AI teams including ServiceNow, IBM, HUMAIN, and Zoho.

Website
https://collinear.ai/
Industry
Software Development
Company size
11-50 employees
Headquarters
Mountain View
Type
Privately Held
Founded
2024
Specialties
LLM, AI, Alignment, Enterprise LLM, RLHF, Red Teaming, AI Evaluation, AI Alignment Infrastructure, Red Teaming for LLMs, Synthetic Data Generation, Model Evaluation Infrastructure, Enterprise AI Deployment, Fine-Tuning and RLHF, LLM Evaluation Loop Automation, Model Improvement Workflows, RL Environments, and Reinforcement Learning

Locations

Employees at Collinear AI

Updates

  • What makes a better training environment for cyber agents: harder tasks, or fairer ones? Either answer goes better with dosa, pani puri, lassi, and a fun crowd of researchers to debate it with! On Thursday, October 1, Collinear AI is bringing researchers and builders together in Sunnyvale for our next Debate on the Frontier and COLM pre-party. Nazneen Rajani will moderate a panel of industry guests from NVIDIA, xAI, MAI, and Google DeepMind. Our debate nights tend to get lively, so come pumped up and ready to join the discussion! After the debate, grab some dosa, pani puri, and mango lassi. Our delicious Collinear cupcakes will be back too. Stick around and hang out with others working on agent evals and training. 📅 October 1, 6:00–8:30 PM 📍 Sunnyvale, California Space is limited. Sign up, and we’ll reserve some cool swag for you. Sign up link in comments.

    • No alternative text description for this image
  • We are on a mission to accelerate AGI for all with frontier data. Cyber is clearly having its moment, but we have been working on it for a while now. Excited to be partnering with Artificial Analysis NVIDIA IBM Vercel on our CWE-bench v1 release 🚀 Read more here: https://cwe-bench.com/

    View organization page for Artificial Analysis

    38,940 followers

    Announcing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with Collinear AI, IBM, NVIDIA, and Vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from Collinear AI covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from Vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized. ➤ CyberGym-E2E-AA, from Berkeley RDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces. Key results: ➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44). ➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%. For more details, read the article: https://lnkd.in/g8t68-mD

    • No alternative text description for this image
  • An agent can identify a vulnerability, write a plausible patch, and still leave the software open to the same exploit. This short video walks through how CWE-bench verifies a fix. A task passes only when the exploit is blocked and the existing regression tests still pass. CWE-bench evaluates defensive coding capabilities across 100 held-out audit-and-patch tasks spanning 54 weakness types. The best model on the current leaderboard reaches 47.8% pass@1. Understanding where an agent failed helps model builders decide what to improve next. Did it miss the vulnerability, leave part of it unfixed, or break working behavior? Collinear’s separate corpus of 1,000+ defensive-security training tasks gives teams a way to work on those gaps. Held-out evaluation tasks are never included in training deliveries. Explore the benchmark and training corpus: https://cwe-bench.com/

  • Google DeepMind used CWE-bench to evaluate Gemini 3.8 Flash Cyber, calling it “a challenging external benchmark for patching capabilities.” 𝐇𝐚𝐫𝐝 𝐛𝐮𝐭 𝐟𝐚𝐢𝐫: no evaluated model clears 50%, and every model faces the same real-world vulnerability-fixing tasks. Read Google DeepMind’s post: Introducing Gemini 3.8 Flash and 3.8 Flash Cyber https://lnkd.in/gKS-_SEK Explore the leaderboard: https://cwe-bench.com/

    • No alternative text description for this image
    • No alternative text description for this image
  • Today we’re launching CWE-bench, a held-out benchmark for the defensive cybersecurity capabilities of frontier coding agents. Real defensive work doesn’t begin with a ticket naming the vulnerability. An agent must audit unfamiliar code, identify what’s wrong, fix every instance, and preserve working behavior. CWE-bench tests that across 100 audit-and-patch tasks built from real open-source projects, spanning 54 weakness types, six programming languages, and all 10 OWASP Top 10 2025 categories. Every task is permanently held out. Agents receive no CVE identifier, file hint, or line number. Memorizing a published patch is not enough: each task introduces additional vulnerabilities of the same class that the upstream fix never covered. A task passes only when the exploit no longer works, and the regression tests still pass. The results show how much work remains:  • Fable 5 leads at 47% pass@1 at maximum reasoning.  • 18 of 100 tasks remain unsolved by every model tested.  • No agent successfully repairs a majority of the benchmark. Explore the benchmark and leaderboard: https://cwe-bench.com/ Read how it was built: https://lnkd.in/gq66u-yT

    • No alternative text description for this image
  • How did we score on the SandboxEscapeBench?

    In the last 4 weeks, we had 2 separate internal cybersecurity incidents involving frontier models. In the first, the agent had to run rollouts, but it didn’t have the API key for one of them. Instead of looking in the researcher’s env, it accessed our cloud-based key storage and picked up a key to run the rollouts. In another incident, the agent ran multiple parallel benchmarks and spun up several EC2 instances. When the agent was done, it not only deleted its own EC2 instances but also another 150 or so odd instances, bringing down several production pipelines. In both instances, the agent initially denied its wrongdoing because the manager agent had no idea how its sub-agents were getting the tasks done. These incidents are only going to get more common as intelligence per watt keeps growing rapidly. Knowing what the agent's attack surface is critical for oversight. One of our research directions is building robust simulated worlds and we believe cybersecurity datasets are well-suited for it. New essay on the topic: https://lnkd.in/g3ftk47g

    • No alternative text description for this image
  • Our researchers were busy at ICML 2026! In between our posters and keynotes, we hosted a dinner with frontier researchers! A few themes we heard across the conference: → Despite coding progress, we noted several benchmarks focused on code. For example, generating time-efficient code was one such focus area. → Interpretability was everywhere. Understanding where LLM failure modes exist is continuing to be key. → Lots of researchers working on RSI, incl. several expressions in spaces like chip design, game design, auto-research and others. Huge thanks to everyone who dropped by to meet us. Easily our favorite night of the conference! Nazneen Rajani Vincent Tu Gonzalo Gonzalez

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Our researchers spent the week at ACL in San Diego, with the sun and some tacos! A few things they came back with: 1. Agent <> simulated-user interaction was everywhere. A lot of people have moved on from one-shot benchmarks. 2. Despite the progress, even coding agents aren't solved. Results are still jagged across tasks within the same domain. 3. Rather than smarter agents, the harder problem now is working out the capability gap frontier. Long but exciting road ahead! Anand Kumar Muyu He Parker Seegmiller

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • We gave 28 AI models $200K and a simulated year to run a startup. The best one 10x'd it. The worst went bankrupt. That's YC-Bench, our long-horizon agent benchmark, and it's one of a few reasons the Collinear team is at ICML in Seoul this week. Where to find us: 🎤 Nazneen Rajani is speaking at the Pluralistic Alignment workshop on July 11. ☕ Our researchers are around all week for coffee chats. Come talk agent evaluation, long-horizon reliability, or simulation. And if you just want to watch your favorite model try to run a company for a year, the YC-Bench leaderboard is live and open-source. Come say hi at ICML, or argue with the leaderboard from home. Both welcome. Links in the comments.

    • No alternative text description for this image

Similar pages

Browse jobs