Introducing the new Lilac
Lilac launched out of YC last August with an open source scheduler and a one line pitch: we find you GPUs. We had watched data scientists burn whole afternoons hunting for free capacity across clouds and on-prem clusters, and we knew one team paying seven million dollars a year to guarantee GPUs that still ran out. Our answer was a scheduler that treated all of it as one pool. Submit the job, and it finds a GPU that fits and runs it.
What finding GPUs taught us
Pretty quickly we noticed something that changed how we thought about the whole market. The hard part wasn't finding GPUs. It was that the GPUs everyone already owned were sitting idle half the time. Enterprise clusters average around 50% utilization, and some on-prem clusters we saw were closer to 10%. The compute already exists. It's just idle.
So we chased that. First as a spot marketplace, reselling interruptible enterprise capacity at spot economics. That taught us a lesson we would keep re-learning: listing idle GPUs is the easy part. Getting anyone to trust hardware they can't see, in a cluster they don't control, is the entire job.
The marketplace became an operator. Providers installed it in their Kubernetes cluster with one kubectl apply, it found reclaimable capacity, ran inference on the idle nodes, and stepped out of the way the second the owner's own jobs came back. On top of that reclaimed capacity we served open weight models (Kimi, GLM, MiniMax's M2.7) at prices that only work when the fixed costs are already paid. It grew fast, at one point 10% week over week, and the model partnerships followed.
The education
Running production inference on other people's clusters will teach you things no amount of talking to customers will. Hardware lies. A node can pass every dashboard check and still quietly miss its numbers. The gap between what we were billed for and what actually ran came out of our margin, so we felt every failure personally. We spent our weeks qualifying capacity, chasing silent failures, and arguing invoices.
To survive it we built tooling: imaging, burn-in, acceptance gates before a node could serve traffic, monitoring that caught the quiet failures, metering accurate enough to dispute a bill with. None of it was the product. All of it was why the product worked.
The pull
Then something started happening that we didn't plan. Providers we worked with saw how we ran their capacity and started asking a different question: can you just manage the whole cluster? Not the idle slice. All of it. Burn-in, monitoring, customer access, billing, support.
The first time, it felt like a distraction. By the third or fourth time it stopped feeling like a distraction and started looking like the actual company. Everyone is buying GPUs. Very few people can run them at the standard the customers (and increasingly the lenders) expect. The layer we kept wishing existed under us, we had accidentally built for ourselves.
What Lilac is now
So now we run GPU fleets for the bare-metal providers who own them. Every node is imaged, burned in, and proven before it bills. We watch it around the clock, wipe it between tenants, swap in validated spares when it fails, and keep one ledger that invoices, SLA credits, and lender audits all settle against. The provider keeps the hardware, the pricing, and the customer. Their customers get a console that behaves like the cloud they were promised.
And if you need GPUs, we still find you GPUs. Reach out and we'll connect you with a partner fleet and get you a quote. If you have GPUs, hand us racks and BMC access and the first nodes usually bill within days.
And the idle problem hasn't gone anywhere. The fleets we manage still have quiet hours, and we still believe that compute should earn. The difference is we finally sit in the right layer to fix it: utilization is an operations problem, and now we run the operations. As the managed platform expands, putting a fleet's idle time back to work is where we're headed, properly this time, from underneath.
Our vision honestly hasn't changed as much as it looks like from the outside. From the scheduler to the marketplace to the operator to the platform, it was always the same belief: the compute already exists, and someone has to make it dependable. We exist to bring stability to the world's AI infrastructure. It just took us a year of running other people's GPUs to understand what that actually required.
Write to us: contact@getlilac.com.
← All news