Local LLM benchmarks and model comparisons, tested on real hardware: an RTX 5090 desktop and a MacBook Pro M5 Max with 36 GB of unified memory. Each video puts the newest open models side by side with measured numbers: accuracy on real tasks, tokens per second, latency, VRAM and peak memory, context length, and how quantization (4-bit, 1-bit, ternary) changes the results. The models covered include Qwen 3.8, Gemma 4, DeepSeek V4, Granite, Nemotron, Bonsai, Laya and the System One decision models like Jev. Every video comes with a full written benchmark on kgptalkie.com with all the tables, so you can check the numbers and pick the right model for your own machine.
📝 Full written benchmarks:
kgptalkie.com/tutorials/llm-benchmarking
🎓 Go deeper:
kgptalkie.com/udemy-courses