Lemony
Reduce LLM costs by up to 90% with intelligent cascade routing
San Francisco, United States · $2.0M raised
- Headquarters
- San Francisco, California
- Employees
- 1–10
- Business Model
- B2B
- Total Funding
- $2.0M
- Last Round
- $2.0M SeedJan 2025
- Rounds
- 1
About
Lemony is an AI infrastructure startup that reduces the cost of running large language models through intelligent cascade routing and domain-specific model selection. Founded by Sascha Buehrle and Ivan Kuleshov and backed by True Ventures with a million seed round, Lemony's core product is cascadeflow, an open-source AI routing system that automatically selects the most cost-efficient model capable of handling each query, achieving reported reductions in monthly LLM spend of up to 90%. The cascadeflow system operates with approximately two milliseconds of routing latency and is designed for both edge devices and cloud deployments. In addition to its cost-optimization router, Lemony has also offered an on-premises generative AI platform that provides enterprise-grade trust, ownership, and transparency in AI by keeping models and data within a company's own infrastructure. The company is headquartered in the San Francisco Bay Area and targets enterprises and developers seeking to make AI usage economically sustainable at scale.
Summary
Lemony is an Artificial Intelligence company based in San Francisco, United States. It has raised $2.0M in total across 1 round, most recently a $2.0M Seed round in Jan 2025. Investors include True Ventures and Alumni Ventures.
Sign up to view the funding chart
Create a free account — you'll get 10 credits a month to unlock funding history, valuations, and tech stacks.
Funding History
Jan 2025
Investors
Similar Companies
Cursor is the best way to code with AI, and built to make you extraordinarily productive.
AI pioneer startup that could make it big in 2026
Provides AI inference hardware and cloud platform for enterprises, using RDUs and SambaCloud to r...

UiPath is an AI-enhanced end-to-end automation platform.
ML inference infrastructure for deploying and serving AI models at scale, performantly and cost-e...

AI hardware company building wafer-scale processors for deep learning training and inference
