vLLM served the same tokens on 4× less energy than Ollama
Open-source AI is consolidating onto hubs you can build on. LiteLLM is the gateway; under it, Ollama and vLLM are both excellent — 7% apart serially, 4.1× apart under load.
Read →Local · Owned · Measured
Open-source AI is consolidating onto hubs you can build on. LiteLLM is the gateway; under it, Ollama and vLLM are both excellent — 7% apart serially, 4.1× apart under load.
Read →The agentic role-playing game I built with Strands Agents at AWS Community Day CEE ran on my own GPUs, fully offline, by the same afternoon — and it is now my standing hybrid AI demo.
A fully self-hosted pipeline: script to voiceover to render, no cloud in the path. Two mascots, worst-to-best, the 3090/4090/DGX numbers, and why the right model depends on your character.
Nvidia is reported to be buying Hugging Face. A good week to learn what is really on the shelf: master weights, vendor builds, and the community quants most of us run.
A year after virtualizing my edge AI supercomputer, I finally measured what the hypervisor costs: single-GPU work is free — the VM even won — and four-way tensor-parallel serving pays a third.
Who is writing this blog: 30 years of datacenters, virtualization and clouds — and now private AI supercomputers, built hands-on. An introduction.