STILLWORKS

LLMSlim v0.3.0

Cut LLM token costs 40-70% with offline prompt compression

Visit product ↗
Developer tools Product Hunt Launched Jul 21, 2026 by Yashvardhan Thanvi yashvardhan_thanvi View original post ↗

v0.3.0 ships hybrid prompt compression: offline TF-IDF extraction pre-prunes context, then an optional LLM rewrite pass semantically optimizes what remains. Result: 40-70% fewer tokens billed, sub-30ms CPU latency, and 100% retention of system directives, code blocks, and JSON schemas. Works with OpenAI, Anthropic, Gemini, LangChain, LlamaIndex, Ollama and more. Zero dependencies. Pure Python. pip install llmslim

Open SourceDeveloper ToolsArtificial IntelligenceGitHub
More Developer tools