LLM Models
All Articles
How to run Bonsai Image, a 1‑bit local image generation model
Bonsai Image is a highly compressed, 1‑bit and 2‑bit image generation model based on Flux that runs fast on consumer GPUs and even Apple Silicon. This guide wal…
Local AI vs trillion‑dollar data centers: how close are we really?
A new wave of local AI models is getting surprisingly close to top cloud models for everyday coding tasks—while costing almost nothing to run. But there’s still…
Kimi K2.6: the open-source coding model with real front-end design taste
Kimi K2.6 is a new open-source AI model focused on front-end design and “vibe coding.” It can generate full Flutter apps, including UI, logic, and even APK buil…
Qwen 3.7 Max: Alibaba’s new flagship model for agents and long-horizon coding
Alibaba’s Qwen 3.7 Max is a new flagship AI model built for agents, long-horizon workflows, and complex coding tasks. It posts frontier-level benchmark scores, …
Can a 35B local model really beat Claude Sonnet 3.5?
Alibaba’s Qwen 3.5 35B model scores higher than Claude Sonnet 3.5 on many benchmarks, but how does it actually perform in real-world coding and front-end tasks?…
How to unlock more VRAM on your Mac for local AI models
Running large language models on a base Apple silicon Mac can feel impossible with limited RAM. This guide explains how macOS handles unified memory, what LM St…
6 Chinese AI models compared: DeepSeek vs Kimi vs Qwen vs GLM vs MiniMax vs MiMo
Six of China’s most advanced large language models were put through three tough, real-world tests: building a production-grade app, reasoning under pressure, an…
Qwen 3.6 Max Preview: Alibaba’s new powerhouse model for coding, agents, and front-end apps
Qwen 3.6 Max Preview is Alibaba’s new flagship model, and it’s already competing with top systems like Claude Opus and Gemini on real-world tasks. From agentic …
ChatGPT GPT 5.5 vs Claude Opus 4.7: real tests, real use cases
Benchmarks don’t tell you which AI model to actually use. This breakdown compares GPT 5.5 and Claude Opus 4.7 across real tasks like research, app building, des…
I tested DeepSeek V4 vs Opus 4.7 vs GPT 5.5: which model should you actually use?
GPT 5.5, Claude Opus 4.7, and DeepSeek V4 all look strong on benchmarks—but how do they behave in real coding workflows? This breakdown compares cost, performan…