Skip to content

Llmash

ollama like local llm interface, optimized to run most models 2-4x faster, even faster than vllm most of the time (tables below) - omgitsbase/llmash

Llmash screenshot

About Llmash

An Ollama-compatible server and command line for Windows, built on llama.cpp.

It serves your GGUF files through llama-server and keeps Ollama's commands, API and model store, so anything already pointed at Ollama keeps working. What it changes is the llama.cpp settings, picked per model at launch.

Nothing needs to be installed first: the installer fetches the llama.cpp build for your GPU, puts llmash on your PATH, and starts the server. With no NVIDIA card it takes the Vulkan or CPU build instead, and models load into system RAM, which works but is slower.

Ollama's store is found and used as it is. A folder of GGUFs from llama.cpp or LM Studio is one command away:

It reads the folder where it is, subfolders included, and says how many models it found. Nothing is copied, downloaded or deleted, and rm will not touch it. llmash models shows what is being read. It is the same setting as OLLAMA_MODELS , which llmash takes the way Ollama does, pointed at either kind of folder.

One GPU, an RTX PRO 6000 Blackwell with 96 GB: same prompts, 4-bit weights in each engine's own format, every backend run as it comes with no hand tuning. Your numbers will differ.

Tokens per second while generating, median of three runs, excluding model load and prompt processing.

The server owns the model store, starts and stops llama-server processes, and decides how long each one stays resident.

Discussion

Sign in to join the discussion.

Loading comments…

Tagged

Alternatives to Llmash

Tools in the same space, ranked by how they are performing in the directory.

Compare all
  • Derive

    Review and approval for work made by AI agents.

    Freemium/Developer Tools
  • Spore

    Run open-weight models on hardware you already own, reach them from any device, and tap a distributed network when you need more.

    Freemium/Developer Tools
  • Conductai

    AI agent governance for teams.

    Open source/Developer Tools
  • Splatit

    An Splatoon server recreation in C++ with a focus on user friendliness and ease of hosting.

    /Developer Tools
  • phpEZ

    Tiny PHP framework. Contribute to QcFe/phpEZ development by creating an account on GitHub.

    /Developer Tools
  • The Simpsons Hit and Run

    Modern Windows version of The Simpsons: Hit & Run, based on the original leaked source code from 2003 - xaskasdf/The-Simpsons-Hit-and-Run

    /Developer Tools