SkillAgentSearch skills...

Gollama

Go manage your Ollama models

Install / Use

npx skills add sammcj/gollama

Installs into whichever agent you are using.

About this skill

Quality Score

0/100

Supported Platforms

Universal

README

Gollama

Gollama is a macOS / Linux tool for managing Ollama models.

It provides a TUI (Text User Interface) for listing, inspecting, deleting, copying, and pushing Ollama models.

The application allows users to interactively select models, sort, filter, edit, run, unload and perform actions on them using hotkeys.

Table of Contents

Features

Gollama is a tool for managing Ollama models with an easy-to-use interface.

It's in active development, so there are some bugs and missing features, however I'm finding it useful for managing my models every day, especially for cleaning up old models.

  • List available models
  • Display metadata such as size, quantisation level, model family, and modified date
  • Edit / update a model's Modelfile
  • Sort models by name, size, modification date, quantisation level, family etc
  • Select and delete models
  • Run and unload models
  • Inspect model for additional details
  • Calculate approximate vRAM usage for a model
  • Copy / rename models
  • Push models to a registry
  • Show running models
  • Has some cool bugs

See also - ingest for passing directories/repos of code to markdown formatted for LLMs.


Update [2025-12-02]: Removal of LM Studio linking & Gollama maintenance slowing

As of the v2.0.1 release of Gollama, LM Studio linking will no longer be available.

Linking from/to LM Studio became more hassle to maintain than it was worth. Ongoing changes to both upstream applications and trying to cater for each users local configuration meant investing too much of my time for a feature I rarely used.

I'm simply not dog-fooding with Ollama enough. This has meant that development has slowed down as I focus on other projects.

I was an early adopter and contributor to Ollama, but the value I got from Ollama has diminished throughout 2025 to the point where I rarely ever use it. For model serving I have mostly moved to llama.cpp running with llama-swap. Llama.cpp has become far more user friendly over the past year, the project is well maintained, easier to configure, with many more features and significantly better performance. For serving models on my laptop I use LM Studio as it provides both MLX models and the standard llama.cpp runtime for GGUF models, in addition to oMLX which has been great for serving MLX models locally for agentic coding with tools like Pi or OpenCode.


Installation

go install (recommended)

go install github.com/sammcj/gollama/v2@latest

curl

I don't recommend this method as it's not as easy to update, but you can use the following command:

curl -sL https://raw.githubusercontent.com/sammcj/gollama/refs/heads/main/scripts/install.sh | bash

Manually

Download the most recent release from the releases page and extract the binary to a directory in your PATH.

e.g. zip -d gollama*.zip -d gollama && mv gollama /usr/local/bin

if "command not found: gollama"

If you see this error, add environment variables to .zshrc or .bashrc.

echo 'export PATH=$PATH:$HOME/go/bin' >> ~/.zshrc
source ~/.zshrc

Usage

To run the gollama application, use the following command:

gollama

Tip: I like to alias gollama to g for quick access:

echo "alias g=gollama" >> ~/.zshrc

Key Bindings

  • Space: Select
  • Enter: Run model (Ollama run)
  • i: Inspect model
  • t: Top (show running models)
  • D: Delete model
  • e: Edit model
  • c: Copy model
  • U: Unload all models
  • p: Pull an existing model
  • ctrl+k: Pull model & preserve user configuration
  • ctrl+p: Pull (get) new model
  • P: Push model
  • n: Sort by name
  • s: Sort by size
  • m: Sort by modified
  • k: Sort by quantisation
  • f: Sort by family
  • B: Sort by parameter size
  • r: Rename model (Work in progress)
  • q: Quit

Top

Top (t)

Inspect

Inspect (i)

Command-line Options

Model Management:

  • -l: List all available Ollama models and exit
  • -s <search term>: Search for models by name
    • OR operator ('term1|term2') returns models that match either term
    • AND operator ('term1&term2') returns models that match both terms
  • -e <model>: Edit the Modelfile for a model
  • -u: Unload all running models
  • -v: Print the version and exit

Configuration:

  • -h, or --host: Specify the host for the Ollama API
  • -H: Shortcut for -h http://localhost:11434 (connect to local Ollama API)
  • --ollama-dir: Custom Ollama models directory
  • --log or --log-level: Override log level (debug, info, warn, error)

Cleanup:

  • --no-cleanup: Don't cleanup broken symlinks

vRAM Analysis:

  • --vram: Estimate vRAM usage for a model. Accepts:
    • Ollama models (e.g. llama3.1:8b-instruct-q6_K, qwen2:14b-q4_0)
    • HuggingFace models (e.g. NousResearch/Hermes-2-Theta-Llama-3-8B)
    • --fits: Available memory in GB for context calculation (e.g. 6 for 6GB)
    • --vram-to-nth or --context: Maximum context length to analyze (e.g. 32k or 128k)
    • --quant: Override quantisation level (e.g. Q4_0, Q5_K_M)
Simple model listing

Gollama can also be called with -l to list models without the TUI.

gollama -l

List (gollama -l):

Edit

Gollama can be called with -e to edit the Modelfile for a model.

gollama -e my-model
Search

Gollama can be called with -s to search for models by name.

gollama -s my-model # returns models that contain 'my-model'

gollama -s 'my-model|my-other-model' # returns models that contain either 'my-model' or 'my-other-model'

gollama -s 'my-model&instruct' # returns models that contain both 'my-model' and 'instruct'
vRAM Estimation

Gollama includes a comprehensive vRAM estimation feature:

  • Calculate vRAM usage for a pulled Ollama model (e.g. my-model:mytag), or huggingface model ID (e.g. author/name)
  • Determine maximum context length for a given vRAM constraint
  • Find the best quantisation setting for a given vRAM and context constraint
  • Shows estimates for different k/v cache quantisation options (fp16, q8_0, q4_0)
  • Automatic detection of available CUDA vRAM (coming soon!) or system RAM

To estimate (v)RAM usage:

gollama --vram llama3.1:8b-instruct-q6_K

📊 VRAM Estimation for Model: llama3.1:8b-instruct-q6_K

| QUANT   | CTX  | BPW | 2K  | 8K              | 16K             | 32K             | 49K             | 64K |
| ------- | ---- | --- | --- | --------------- | --------------- | --------------- | --------------- |
| IQ1_S   | 1.56 | 2.2 | 2.8 | 3.7(3.7,3.7)    | 5.5(5.5,5.5)    | 7.3(7.3,7.3)    | 9.1(9.1,9.1)    |
| IQ2_XXS | 2.06 | 2.6 | 3.3 | 4.3(4.3,4.3)    | 6.1(6.1,6.1)    | 7.9(7.9,7.9)    | 9.8(9.8,9.8)    |
| IQ2_XS  | 2.31 | 2.9 | 3.6 | 4.5(4.5,4.5)    | 6.4(6.4,6.4)    | 8.2(8.2,8.2)    | 10.1(10.1,10.1) |
| IQ2_S   | 2.50 | 3.1 | 3.8 | 4.7(4.7,4.7)    | 6.6(6.6,6.6)    | 8.5(8.5,8.5)    | 10.4(10.4,10.4) |
| IQ2_M   | 2.70 | 3.2 | 4.0 | 4.9(4.9,4.9)    | 6.8(6.8,6.8)    | 8.7(8.7,8.7)    | 10.6(10.6,10.6) |
| IQ3_XXS | 3.06 | 3.6 | 4.3 | 5.3(5.3,5.3)    | 7.2(7.2,7.2)    | 9.2(9.2,9.2)    | 11.1(11.1,11.1) |
| IQ3_XS  | 3.30 | 3.8 | 4.5 | 5.5(5.5,5.5)    | 7.5(7.5,7.5)    | 9.5(9.5,9.5)    | 11.4(11.4,11.4) |
| Q2_K    | 3.35 | 3.9 | 4.6 | 5.6(5.6,5.6)    | 7.6(7.6,7.6)    | 9.5(9.5,9.5)    | 11.5(11.5,11.5) |
| Q3_K_S  | 3.50 | 4.0 | 4.8 | 5.7(5.7,5.7)    | 7.7(7.7,7.7)    | 9.7(9.7,9.7)    | 11.7(11.7,11.7) |
| IQ3_S   | 3.50 | 4.0 | 4.8 | 5.7(5.7,5.7)    | 7.7(7.7,7.7)    | 9.7(9.7,9.7)    | 11.7(11.7,11.7) |
| IQ3_M   | 3.70 | 4.2 | 5.0 | 6.0(6.0,6.0)    | 8.0(8.0,8.0)    | 9.9(9.9,9.9)    | 12.0(12.0,12.0) |
| Q3_K_M  | 3.91 | 4.4 | 5.2 | 6.2(6.2,6.2)    | 8.2(8.2,8.2)    | 10.2(10.2,10.2) | 12.2(12.2,12.2) |
| IQ4_XS  | 4.25 | 4.7 | 5.5 | 6.5(6.5,6.5)    | 8.6(8.6,8.6)    | 10.6(10.6,10.6) | 12.7(12.7,12.7) |
| Q3_K_L  | 4.27 | 4.7 | 5.5 | 6.5(6.5,6.5)    | 8.6(8.6,8.6)    | 10.7(10.7,10.7) | 12.7(12.7,12.7) |
| IQ4_NL  | 4.50 | 5.0 | 5.7 | 6.8(6.8,6.8)    | 8.9(8.9,8.9)    | 10.9(10.9,10.9) | 13.0(13.0,13.0) |
| Q4_0    | 4.55 | 5.0 | 5.8 | 6.8(6.8,6.8)    | 8.9(8.9,8.9)    | 11.0(11.0,11.0) | 13.1(13.1,13.1) |
| Q4_K_S  | 4.58 | 5.0 | 5.8 | 6.9(6.9,6.9)    | 8.9(8.9,8.9)    | 11.0(11.0,11.0) | 13.1(13.1,13.1) |
| Q4_K_M  | 4.85 | 5.3 | 6.1 | 7.1(7.1,7.1)    | 9.2(9.2,9.2)    | 11.4(11.4,11.4) | 13.5(13.5,13.5) |
| Q4_K_L  | 4.90 | 5.3 | 6.1 | 7.2(7.2,7.2)    | 9.3(9.3,9.3)    | 11.4(11.4,11.4) | 13.6(13.6,13.6) |
| Q5_K_S  | 5.54 | 5.9 | 6.8 | 7.8(7.8,7.8)    | 10.0(10.0,10.0) | 12.2(12.2,12.2) | 14.4(14.4,14.4) |
| Q5_0    | 5.54 | 5.9 | 6.8 | 7.8(7.8,7.8)    | 10.0(10.0,10.0) | 12.2(12.2,12.2) | 14.4(14.4,14.4) |
| Q5_K_M  | 5.69 | 6.1 | 6.9 | 8.0(8.0,8.0)    | 10.2(10.2,10.2) | 12.4(12.4,12.4) | 14.6(14.6,14.6) |
| Q5_K_L  | 5.75 | 6.1 | 7.0

Related Skills

View on GitHub
GitHub Stars1.8k
CategoryDevelopment
Updated2h ago
Forks109

Languages

Go

Security Score

100/100

Audited on Aug 8, 2026

No findings