Gotchi
All posts
BenchmarksApril 20, 20265 min read

DeepSeek V3: local benchmark results

We tested DeepSeek V3 against Llama 3.1 70B on Apple Silicon. Results were surprising - especially on reasoning tasks.

Gotchi Team
Engineering

DeepSeek V3 landed last month and the hype was real. We ran comprehensive benchmarks on Apple Silicon hardware to see how it performs locally compared to the reigning champion, Llama 3.1 70B.

Test setup

All tests ran on a MacBook Pro M4 Pro with 48GB unified memory. Both models used Q4_K_M quantization. We measured tok/s, MMLU-Pro accuracy, HumanEval pass@1, and real-world task completion across email summarization, code generation, and research synthesis.

Key findings

  • DeepSeek V3 scored 89.4% on coding tasks vs Llama's 86.2% - a meaningful 3.2-point advantage.
  • On reasoning and research, DeepSeek led by 4.2 points (92.1% vs 87.9%).
  • Llama 3.1 was faster: 340ms average latency vs DeepSeek's 280ms (DeepSeek's architecture is more efficient per parameter).
  • Both models handled email summarization equally well - 95%+ accuracy on our internal benchmark.
  • DeepSeek V3 uses 32GB disk space vs Llama's 40GB in Q4_K_M, making it more accessible on 32GB Macs.

DeepSeek V3 is the first open-source model that makes me forget I'm not using Claude.

- Internal tester
Try it now

DeepSeek V3 is available in Gotchi's model library. One click to download, ready to chat in minutes.

Ready to run AI locally?

Download Gotchi. Open core, local-first, and ready for every model you want to run.