Sail Research’s cover photo
Sail Research

Sail Research

Software Development

San Francisco, California 2,687 followers

Infrastructure for long-horizon agents.

About us

Infrastructure for long-horizon agents. Efficient inference and sandboxes optimized for throughput. 10x more cost efficient for your token budgets.

Website
https://sailresearch.com
Industry
Software Development
Company size
11-50 employees
Headquarters
San Francisco, California
Type
Privately Held

Locations

  • Primary

    600 California Street

    San Francisco, California 94108, US

    Get directions

Employees at Sail Research

Updates

  • 10× lower AI cost per product. Channel3 is building the product graph for agentic commerce, which takes trillions of tokens of inference to keep current. But not every request needs an instant answer. The team moved the latency-tolerant stages of their product graph pipeline to asynchronous inference with Gemma 4 31B on Sail's Flex completion window. Background jobs finish with a p95 of 8.5 minutes, and the API still returns products in under half a second with the same output quality. "Sail lets us use latency as a lever and run dramatically more AI for the same budget," says Alex Schiff, CEO of Channel3. This is exactly the workload we built Sail for. Channel3 adds millions of products every month, and we're excited to keep building together as they bring every product on the internet within reach of AI agents. Read more at the link in the comments!

    • No alternative text description for this image
  • We've assembled an amazing team and are working on massive problems. Give our careers page a visit :) https://lnkd.in/gxJks93p

    After 2008, Nvidia cut back on perks. They got rid of free lunch and even free milk. If you wanted milk in your coffee, you had to pay $1 a month to join the “milk club,” which stocked the fridge with Costco milk. Neil Movva remembers it as one small example of the frugality that permeated Nvidia when he worked there on GPUs and kernels. Neil has an unusually deep understanding of inference, from software to chips to power. We spend a lot of time on each of those layers, how they connect, and where the important tradeoffs are. What makes this conversation special is how detailed it is (like a 401-level class), yet Neil makes it remarkably clear and easy to follow. Today he runs Sail Research, a company building infrastructure for agents to make tokens as cheap as possible. We discuss latency versus throughput, why there are no bad chips only bad pricing, buying chips and power nobody else wants, Nvidia, open source, and the economics of the frontier labs. I learned a ton. Enjoy! https://lnkd.in/gp4gZvDn

  • Sail Research reposted this

    After 2008, Nvidia cut back on perks. They got rid of free lunch and even free milk. If you wanted milk in your coffee, you had to pay $1 a month to join the “milk club,” which stocked the fridge with Costco milk. Neil Movva remembers it as one small example of the frugality that permeated Nvidia when he worked there on GPUs and kernels. Neil has an unusually deep understanding of inference, from software to chips to power. We spend a lot of time on each of those layers, how they connect, and where the important tradeoffs are. What makes this conversation special is how detailed it is (like a 401-level class), yet Neil makes it remarkably clear and easy to follow. Today he runs Sail Research, a company building infrastructure for agents to make tokens as cheap as possible. We discuss latency versus throughput, why there are no bad chips only bad pricing, buying chips and power nobody else wants, Nvidia, open source, and the economics of the frontier labs. I learned a ton. Enjoy! https://lnkd.in/gp4gZvDn

  • Sail Research reposted this

    Today we're launching Sailboxes: the first cloud environment purpose-built for long-horizon AI agents. Sailboxes are full machines with persistent state that auto-sleep when idle, so waiting on input, inference, or results costs nothing. No runtime limits means agents can run for days or weeks. Need to scale? Fork a running Sailbox instead of provisioning from scratch. All from $0.015 per active vCPU-hour, over 70% cheaper than other providers. Under the hood: novel architecture from Nirvik Baruah and Charley Cunningham that live-migrates VMs based on actual resource usage. To prove how robust it is, we ran a Minecraft server in a Sailbox and forced it to migrate every 2 minutes--players never noticed. Each Sailbox is a full machine with independent disk, Docker support, local NVMe, persistent state. Memory is elastic too: boxes start small, grow up to 128 GB as your workload needs, and shrink back automatically. Your agents behave the same in the cloud as they do locally. Thousands are already in production: hosting agents, running RL rollouts, powering coding + triage workflows. We’re excited to partner with our friends Quadrillion Labs who are already running parallel research experiments on Sailboxes via Qualia Cloud. Spin up your Sailbox today at the link below in comments!

  • Today we're launching Sailboxes: the first cloud environment purpose-built for long-horizon AI agents. Sailboxes are full machines with persistent state that auto-sleep when idle, so waiting on input, inference, or results costs nothing. No runtime limits means agents can run for days or weeks. Need to scale? Fork a running Sailbox instead of provisioning from scratch. All from $0.015 per active vCPU-hour, over 70% cheaper than other providers. Under the hood: novel architecture from Nirvik Baruah and Charley Cunningham that live-migrates VMs based on actual resource usage. To prove how robust it is, we ran a Minecraft server in a Sailbox and forced it to migrate every 2 minutes--players never noticed. Each Sailbox is a full machine with independent disk, Docker support, local NVMe, persistent state. Memory is elastic too: boxes start small, grow up to 128 GB as your workload needs, and shrink back automatically. Your agents behave the same in the cloud as they do locally. Thousands are already in production: hosting agents, running RL rollouts, powering coding + triage workflows. We’re excited to partner with our friends Quadrillion Labs who are already running parallel research experiments on Sailboxes via Qualia Cloud. Spin up your Sailbox today at the link below in comments!

  • And we're Sailing! We're building a future with abundant, efficient inference. Stay tuned, we're excited to share more announcements soon

    View profile for Neil Movva

    Samir Menon and I are thrilled to announce Sail Research! We build infrastructure for long-horizon agents: inference served at unbeatable price-per-token for open models, plus sandboxes designed to run for days, weeks, or even longer. We've raised $80M, with our seed led by Sequoia Capital and series A led by Kleiner Perkins. What makes agents so different? Instead of racing to serve a human waiting at a keyboard, agents need scale, reliability, and sustainable cost. Sail finds this efficiency everywhere in the stack: we carefully choose our chips, write custom inference engines, and run a global controller that fully utilizes every computer in our fleet. Tight integration from silicon to API lets Sail open up the cost / latency frontier to our customers - the most patient agents can now access 10x more intelligence per dollar. We're excited to be working with great companies like Parallel Web Systems, Detail, Jack & Jill, and Quadrillion Labs to deploy long-horizon agents with trillions of tokens. Our team is thoughtful in our engineering craft and relentlessly ambitious in our pursuit of peak performance. We previously trained at companies like NVIDIA, OpenAI, Google, and so many trading firms. Now we're ready to do the work that will define our careers, in the most compute intensive market of all time. Welcome to the era of abundant intelligence. We can't wait to build with you!

Similar pages