Skip to content
GHMyGearHut
[EMPIRICAL HARDWARE LABS]·Physical Benchmarks & Local Compute

Hardware Stress Tests & Labs

Did we actually test this? Physical workstations, local Ollama/vLLM throughput, memory pressure logs, and reproducible benchmarks.

LAB REPORT 01Hardware Benchmark2026-09-07

Lab Report: Claude Code vs. Cursor Agent on 50 Full-Stack Refactors

We benchmarked Claude Code CLI and Cursor Agent across 50 production repository migrations. Here are the hard metrics on token cost, execution time, and error rates.

TEST CONCLUSION

Claude Code CLI completed 44/50 refactors without syntax breakage due to subagent task isolation, while Cursor Agent completed edits 2.4x faster on single-file modifications.

Reproducible Data
Read Methodology
LAB REPORT 02Hardware Benchmark2026-09-01

Lab Report: Benchmarking Qwen 2.5 Coder 32B on Apple Mac Studio M3 Max

In-depth testing of token generation throughput, memory pressure, and 32k context retention for local AI coding assistants.

TEST CONCLUSION

Apple Silicon unified memory delivers 32 tokens/second on 32B quantized models with zero fan noise, proving viable for 100% private developer environments.

Reproducible Data
Read Methodology