We combine local models and hosted ones to produce cheap tokens that rival the frontier.
imp is a new kind of inference provider. Our engine squeezes the most out of your personal machine to run capable local models efficiently, and balances them out with larger, more capable hosted models to perform tasks at a fraction of the cost of frontier models.
imp comes as a lightweight daemon that your existing AI tools and applications can tap into, making them faster and absurdly cheaper to use without compromises. It runs on high-end machines equipped with Apple Silicon (with support for other unified memory systems coming after), and selects optimal models based on your hardware.
We believe that balancing out on-device and hosted inference is an effective way to drive down the cost of intelligence and enable a new class of applications to be built. If this is something you'd like to work on, reach out.
imp is currently in development and will publicly launch later this year.