Gotchi
Conversational AI

Your AI, running on your hardware

Chat with the most capable open-source models - locally. No API keys, no usage caps, no data leaving your machine.

Capabilities

Everything you expect from AI, nothing in the cloud

Gotchi brings frontier-class capabilities to your desktop without sending your data to the cloud.

Natural conversation

Multi-turn chat with full context windows. Ask follow-ups, refine responses, and build on previous threads - just like a cloud chatbot, but private.

Multi-model support

Switch between Llama, DeepSeek, Qwen, Mixtral, Gemma, Phi, and the rest of the open library. Run the model that fits your task, not the vendor's default.

Tool use & function calling

Models can browse the web, execute shell commands, read files, and call APIs. Gotchi handles tool routing so models act, not just talk.

Advanced reasoning

Chain-of-thought, step-by-step breakdowns, and structured output. Let the model think before it answers for better accuracy on complex problems.

Zero data exposure

Every inference runs on your CPU or GPU. Your prompts, documents, and conversations never leave your hardware - period.

Fine-grained controls

Adjust temperature, top-p, max tokens, system prompts, and stop sequences per conversation. Full control over how your AI responds.

Local inference

No internet required. Seriously.

Gotchi uses Ollama under the hood to run quantized models directly on your machine. Whether you're on a plane, in a coffee shop, or behind a corporate firewall - your AI works. Models run on Apple Silicon, NVIDIA GPUs, or plain CPU. No fallback to the cloud, ever.
Local inferenceOffline
Your prompt
Model reasoning
GPU compute
Knowledge base
Everything stays on your machine
Benchmarks

Real performance on real hardware

Token generation speed varies by model size and hardware. Here are representative numbers from Apple Silicon Macs.

Phi-3.5 3.8B
92 tok/s
M2 Air·8GB
Llama 3.1 8B
54 tok/s
M3 Pro·18GB
Qwen 2.5 32B
28 tok/s
M4 Pro·48GB
Llama 3.1 70B
14 tok/s
M4 Max·128GB
Context & memory

AI that remembers what matters

Gotchi builds a personal knowledge graph from your conversations, indexed documents, and connected apps. Your AI learns your preferences, projects, and patterns over time - all stored locally in a private vector database that only you can access.
Memory
Search your knowledge base
340Chats
128Docs
12Repos
"What did we decide about rate limiting?"
Found in Slack, RFC-042, meeting notes
Built for power users

Go beyond chat

Knowledge base

Index PDFs, markdown, code repos, and notes. Ask questions grounded in your own data with retrieval-augmented generation.

Command palette

Trigger AI from anywhere on your desktop with a global shortcut. Summarize selected text, rewrite emails, or generate code without switching apps.

Start a conversation with your own AI

Download the free Gotchi core and run your first model in under two minutes.

macOS·Windows·No telemetry·Free core · €49.99/yr all-in · Free updates