Gotchi
All posts
EngineeringApril 28, 20266 min read

Why we built on Ollama

Ollama handles model management beautifully. We decided to build on it rather than reinvent the wheel. Here's why.

Gotchi Bloger
Editorial

Every startup faces build-vs-buy decisions. For local model inference - the core of everything Gotchi does - we chose to build on Ollama rather than writing our own runtime. It's one of the best technical decisions we've made.

What Ollama does well

Ollama handles the hardest parts of local inference: memory mapping, GPU offloading, context window management, and quantization. It supports every major model architecture, gets new model support within days of release, and runs on macOS, Linux, and Windows.

  • Battle-tested runtime with millions of users
  • Metal acceleration on Apple Silicon out of the box
  • NVIDIA CUDA support for GPU offloading
  • Automatic memory management for large models
  • New model support within days of release

What Gotchi adds

Ollama is the engine. Gotchi is the car. We wrap Ollama with a native desktop app, agent workflows, connectors, a knowledge base, personal memory, local usage insights, and privacy controls. Ollama users get everything they wished the CLI had.

We didn't build Gotchi because Ollama was bad. We built it because Ollama was good - and deserved a better interface.

The decision framework

If you're building a product on local AI, ask yourself: is inference our differentiator? For us, the answer was no - our value is in the UX, agents, and integrations layer. Ollama let us focus on what we do best while shipping a better product faster.

Ready to run AI locally?

Download Gotchi. Open core, local-first, and ready for every model you want to run.