A seamless middleware that compresses your LLM prompts. Slash your API bills by up to 35% without degrading model performance.
npm i promptyWorks with OpenAI, Anthropic & Gemini
Works with every LLM provider
Built for scale.
A high-performance text compression engine that processes prompts in under 20ms. Completely transparent to your users.
Zero Latency
Runs entirely on your server or at the edge. No external API calls required for the core compression engine.
Universal Compatibility
Outputs standard text. If your LLM accepts strings, it works. First-class support for all major providers.
Cryptographic Privacy
Your prompts never leave your infrastructure. Prompty can be deployed in a Trusted Execution Environment (TEE).
Integrate in minutes,
save forever.
Our drop-in SDK automatically intercepts and compresses outgoing prompts before they hit the LLM provider. Two lines of code is all it takes.
- ✓Typed for TypeScript
- ✓Next.js App Router ready
- ✓Edge compatible
import { compress } from 'prompty';import OpenAI from 'openai';// 1. Wrap your promptconst { compressed } = compress(systemPrompt);// 2. Call the LLM normallyconst openai = new OpenAI();const res = await openai.chat.completions.create({model: 'gpt-4-turbo',messages: [{ role: 'system', content: compressed }],});Ready to shrink
your AWS bill?
I built this because LLM token costs were killing my side projects. Now it's open source and free for everyone. Give it a try.