Gotchi
Local model library

Every open model, one click to run

Browse the full open-source model catalog, download what you need, and run it on your hardware - no cloud account required.

Model management

Every model, managed for you

Gotchi handles downloads, updates, quantization, and hardware detection so you can focus on using AI, not configuring it.

One-click install

Pick a model from the library and Gotchi pulls it down, verifies the checksum, and loads it - ready to chat in seconds.

Hardware-aware loading

Gotchi detects your GPU, VRAM, and unified memory to recommend the best quantization level. Apple Silicon, NVIDIA, and CPU-only all supported.

Quantization options

Choose between Q4, Q5, Q6, Q8, and full-precision weights. Smaller quants run faster on limited hardware with minimal quality loss.

Auto-updates

When a model publisher releases a new version, Gotchi notifies you and offers a one-click upgrade - no manual downloading or swapping.

Local storage control

See exactly how much disk each model uses. Delete, archive, or move models between drives without breaking your setup.

Performance benchmarks

Built-in tok/s benchmarking for every model on your hardware. Compare models side by side before committing to a workflow.

Powered by Ollama

Built on a battle-tested engine

Gotchi wraps Ollama with a polished interface, automatic lifecycle management, and seamless model switching. Ollama's battle-tested runtime handles memory mapping, GPU offloading, and context management - Gotchi handles everything else.
Ollama runtime
Model weights40 GB
Metal accelerationM4 Pro
Inference42 tok/s
Llama 3.1 70B ready
Hardware

Optimized for every chip

Gotchi auto-detects your hardware and recommends the best quantization. Apple Silicon M1–M4 chips get Metal acceleration. NVIDIA GPUs use CUDA offloading. Even CPU-only machines can run 7B models comfortably. More VRAM means bigger models - it's that simple.
Hardware
Apple SiliconMetal
NVIDIA GPUCUDA
CPU onlyAVX2
AMD GPUsoon
Detected: Apple M4 Pro · 48 GB
Supported models

If it runs on Ollama, it runs on Gotchi

New models are available the day they launch. Here are some of the most popular families.

Llama 3.1 · 3.2 · 3.3
DeepSeek-R1 · V3 · Coder
Qwen 2.5 · QwQ · Coder
Mixtral 8x7B · 8x22B
Gemma 2 · 3
Phi-3 · Phi-4
Mistral · Nemo · Large
CodeLlama · StarCoder2
Command R · R+
Yi · InternLM · GLM
Vicuna · OpenHermes
LLaVA · BakLLaVA (vision)

Run any size, on any machine

Multi-model workflows

Load multiple models simultaneously and route tasks to the best one. Use a small model for quick answers and a large one for deep analysis.

Optimized for your hardware

Metal acceleration on Apple Silicon and CUDA/AVX2 on Windows give near-native speed for 7B–70B parameter models. Gotchi detects your hardware and picks the right build automatically.

Your model library awaits

Download Gotchi and start exploring the open-source model library - local, private, and always available.

macOS·Windows·No telemetry·Free core · €49.99/yr all-in · Free updates