AI

Run and serve open large language models

Run open large language models on remote compute — served to your browser with a chat interface and an HTTP API for your own apps.

This category covers:

  • Local LLMs — run open models (Llama, Qwen, Gemma, and more) with Ollama, GPU-accelerated when available

Browse the workflows below. To deploy one, see How it works.