Picolm
SIMD/GPU pure C inference for Llama 2-family, GPT-2, Gemma-3n, and Qwen 3.x GGUF files.
About Picolm
Note: MoE model GPU tensor upload is not yet supported on HIP/ROCm (CUDA works fine). MoE models on HIP fall back to CPU SSM matmul; full GPU upload path planned for next release.
MSYS2 CUDA build target: hunger-msys-direct Makefile target for MSYS2 SSH builds with nvcc+cl.exe
The original PicoLM was a clever proof-of-concept: mmap a GGUF, stream layers through 45MB of RAM, call it a day. Noble effort, but it was missing basically everything you need to actually use an LLM. So I went in and added hundreds of commits across quantization acceleration, model support, HTTP server, GPU backends, and cross-platform fixes.
Since then, progress has been steady. The main goals are (beyond satisfying curiosity on what Qwen-3.6-27B can and can't do):
First, portability. The only good software is one that runs (and is tested) on everything from the last 50 years. PicoLM runs from 32-bit MS-DOS through Raspberry Pi, MIPS/OpenWRT, AMD ROCm (MI50 tested), Metal (if someone bothers), to RTX 4090 or DGX Spark.
Discussion
Sign in to join the discussion.
Loading comments…
Tagged
Alternatives to Picolm
Tools in the same space, ranked by how they are performing in the directory.
Supabase
Build production-grade applications with a Postgres database, Authentication, instant APIs, Realtime, Functions, Storage and Vector embeddings.
Open source/Developer ToolsHillock
Local, gradient-free neuro-symbolic memory engine combining Hyperdimensional Computing (HDC/VSA), Hebbian plasticity, and graph triples for offline AI.
/Developer ToolsShaide
Distributed, multi-model LLM inference on Kubernetes you own - axem-solutions/shaide
Open source/InfrastructureWeb LLM
High-performance In-browser LLM Inference Engine .
Open source/InfrastructureAgentic Data Kernel
Temporal knowledge and durable workflow infrastructure for software agents - Jason-Doyle/agentic-data-kernel
Open source/Infrastructure- /Developer Tools