Llama.cpp
Portable local inference engine for running quantized language models across laptops, desktops, and servers. It is infrastructure rather than a hosted ChatGPT replacement, so you supply the model and hardware.
Choose a paid tool and browse free or open replacements for the same job.
Browse by job
Start with a category, or search the full catalog below.
All results
Search, filter, and choose a useful replacement.
Portable local inference engine for running quantized language models across laptops, desktops, and servers. It is infrastructure rather than a hosted ChatGPT replacement, so you supply the model and hardware.
Enhanced ChatGPT clone with multi-provider support.
Chat UI that works with any LLM. It comes loaded with advanced features like agents, web search, RAG, MCP, deep research, Connectors to 40+ knowledge sources, and more.
User-friendly AI Interface, supports Ollama, OpenAI API.
Meta · 8B-405B · Best open-weight model family
Mistral AI · 7B-8x22B · Fast, efficient, MoE architecture
Alibaba · 0.5B-72B · Strong multilingual + coding
Google · 2B-27B · Lightweight, efficient