Source profileQuality 92/100Review permissions

Jamie-BitFlight/claude_skills/plugins/llamafile/skills/llamafile/SKILL.md

llamafile

When setting up local LLM inference without cloud APIs. When running GGUF models locally. When needing OpenAI-compatible API from a local model. When building offline/air-gapped AI tools. When troubleshooting local LLM server connections.

Source repository stars
64
Declared platforms
0
Static risk flags
3
Last source update
2026-08-28
Source checked
2026-08-28

Decision brief

What it does: where it fits

Configure and manage Mozilla Llamafile - a cross-platform executable distribution format that runs LLMs locally with an OpenAI-compatible API.

Best for

  • Installing llamafile binary and GGUF model files
  • Starting llamafile server with optimal configuration
  • Integrating llamafile with LiteLLM or OpenAI SDK

Not for

  • Server Fails to Start
  • Connection Refused Errors

Compatibility matrix

Platform support, with evidence labels

PlatformStatusEvidenceWhat to check
CodexNot declaredNo explicit evidencePortability before use
Claude CodeNot declaredNo explicit evidencePortability before use
CursorNot declaredNo explicit evidencePortability before use
Gemini CLINot declaredNo explicit evidencePortability before use
Open the compatibility checker

Installation

Inspect first. Install second.

The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

Source-detected install commandSource
npx skills add https://github.com/Jamie-BitFlight/claude_skills --skill "plugins/llamafile/skills/llamafile"
Safe inspection promptEditorial

Inspect the Agent Skill "llamafile" from https://github.com/Jamie-BitFlight/claude_skills/blob/a00194f25fec502d3d659b7d610369614967251e/plugins/llamafile/skills/llamafile/SKILL.md at commit a00194f25fec502d3d659b7d610369614967251e. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

Workflow

What the source asks the agent to do

  1. 01

    Process Management Script

    Python script to start llamafile as background process with health checking:

    Python script to start llamafile as background process with health checking:
  2. 02

    Find process using port 8080

    Review the “Find process using port 8080” section in the pinned source before continuing.

    Review and apply the “Find process using port 8080” source section.
  3. 03

    Kill existing process

    kill $(lsof -t -i :8080) bash ls -lh /path/to/model.gguf bash ls -la /path/to/llamafile

    kill $(lsof -t -i :8080) bash ls -lh /path/to/model.gguf bash ls -la /path/to/llamafile
  4. 04

    When to Use This Skill

    Installing llamafile binary and GGUF model files

    Installing llamafile binary and GGUF model filesStarting llamafile server with optimal configurationIntegrating llamafile with LiteLLM or OpenAI SDK
  5. 05

    Core Capabilities

    Llamafile combines llama.cpp with Cosmopolitan Libc to create single-file executables that:

    Run on macOS, Windows, Linux, FreeBSD, OpenBSD, NetBSDSupport AMD64 and ARM64 architecturesServe OpenAI-compatible HTTP API on localhost

Permission review

Static risk signals and limitations

Writes files

medium · line 25

The documentation asks the agent to create, modify, or delete local files.

Llamafile combines llama.cpp with Cosmopolitan Libc to create single-file executables that:

Network access

medium · line 54

The documentation includes network, browsing, or remote request actions.

curl -L -o llamafile https://github.com/mozilla-ai/llamafile/releases/download/0.9.3/llamafile-0.9.3

Runs scripts

medium · line 56

The documentation asks the agent to run terminal commands or scripts.

# Make executable

Network access

medium · line 76

The documentation includes network, browsing, or remote request actions.

curl -LO https://huggingface.co/mozilla-ai/llava-v1.5-7b-llamafile/resolve/main/llava-v1.5-7b-q4.llamafile

Evidence record

Why each signal appears

EvidenceSourceComputedTestedEditorial
SignalValueEvidence typeMeaning
Quality score92/100ComputedDocumentation, specificity, maintenance, and trust rules
Repository stars64SourceRepository attention, not individual Skill quality
Compatibility0 platformsSourceDeclared in the catalog source record
Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

Pinned source

Provenance and original SKILL.md

Repository
Jamie-BitFlight/claude_skills
Skill path
plugins/llamafile/skills/llamafile/SKILL.md
Commit
a00194f25fec502d3d659b7d610369614967251e
License
MIT
Collected
2026-08-28
Default branch
main
View the original SKILL.md

Llamafile

Configure and manage Mozilla Llamafile - a cross-platform executable distribution format that runs LLMs locally with an OpenAI-compatible API.

When to Use This Skill

Use this skill when:

  • Installing llamafile binary and GGUF model files
  • Starting llamafile server with optimal configuration
  • Integrating llamafile with LiteLLM or OpenAI SDK
  • Configuring llamafile for different performance profiles (GPU, CPU, network access)
  • Troubleshooting llamafile server startup or API connection issues
  • Building applications requiring local LLM inference
  • Setting up commit message tools, code review systems, or other developer tools with local AI
  • Managing llamafile as a background service
  • Selecting and downloading appropriate GGUF models
  • Validating OpenAI-compatible API responses

Core Capabilities

What Llamafile Provides

Llamafile combines llama.cpp with Cosmopolitan Libc to create single-file executables that:

  • Run on macOS, Windows, Linux, FreeBSD, OpenBSD, NetBSD
  • Support AMD64 and ARM64 architectures
  • Serve OpenAI-compatible HTTP API on localhost
  • Load GGUF model files for inference
  • Provide /health endpoint for monitoring
  • Support GPU acceleration (CUDA, Metal, Vulkan)
  • Enable embeddings generation with --embedding flag

API Compatibility

Llamafile exposes these OpenAI-compatible endpoints when running with --server:

EndpointDescriptionRequirements
http://localhost:8080/v1/chat/completionsChat completions (primary)Server mode
http://localhost:8080/v1/completionsText completionsServer mode
http://localhost:8080/v1/embeddingsGenerate embeddings--embedding flag
http://localhost:8080/healthHealth checkServer mode

Critical Detail: All OpenAI-compatible endpoints require /v1 prefix in the URL path.

Installation

Download Llamafile Binary

# Download llamafile v0.9.3 binary
curl -L -o llamafile https://github.com/mozilla-ai/llamafile/releases/download/0.9.3/llamafile-0.9.3

# Make executable
chmod 755 llamafile

# Verify version
./llamafile --version

Alternative download sources:

  • GitHub Release: https://github.com/mozilla-ai/llamafile/releases/download/0.9.3/llamafile-0.9.3
  • SourceForge Mirror: https://sourceforge.net/projects/llamafile.mirror/files/0.9.3/

Download a Model

Llamafile supports two approaches: pre-packaged llamafile executables (model embedded) or separate GGUF model files.

Pre-packaged llamafile (easiest):

# Download a llamafile with embedded model
curl -LO https://huggingface.co/mozilla-ai/llava-v1.5-7b-llamafile/resolve/main/llava-v1.5-7b-q4.llamafile
chmod +x llava-v1.5-7b-q4.llamafile
./llava-v1.5-7b-q4.llamafile --server --nobrowser

Separate GGUF model (use with ./llamafile --server -m model.gguf):

Download GGUF files from HuggingFace model publishers, then load with the llamafile binary.

Pre-packaged llamafile models from mozilla-ai:

ModelSizeUse CaseDownload
Qwen3-0.6B~500MBFast, lower qualitymozilla-ai/Qwen3-0.6B-llamafile
Mistral 7B v0.2~4GBBalanced speed/qualitymozilla-ai/Mistral-7B-Instruct-v0.2-llamafile
Llama 3.1 8B~5GBHigher quality, slowermozilla-ai/Meta-Llama-3.1-8B-Instruct-llamafile
LLaVA v1.5 7B~4GBMultimodal (text+image)mozilla-ai/llava-v1.5-7b-llamafile

These are self-contained executables — download, chmod +x, and run. No separate llamafile binary needed.

Server Configuration

Basic Server Command

Start llamafile server for local API access:

./llamafile --server \
    -m /path/to/model.gguf \
    --nobrowser \
    --port 8080 \
    --host 127.0.0.1

Critical flags explained:

  • --server: Required to enable HTTP API endpoints
  • -m: Path to GGUF model file (required)
  • --nobrowser: Prevents auto-opening browser on startup
  • --port 8080: Default port (note: NOT 8000)
  • --host 127.0.0.1: Localhost only (secure default)

Performance-Optimized Configuration

For GPU-accelerated inference with higher throughput:

./llamafile --server \
    -m /path/to/model.gguf \
    --nobrowser \
    --port 8080 \
    --host 127.0.0.1 \
    --ctx-size 4096 \
    --n-gpu-layers 99 \
    --threads 8 \
    --cont-batching \
    --parallel 4

Advanced flags:

FlagPurposeDefaultWhen to Use
--ctx-sizePrompt context window size512Increase for longer conversations
--n-gpu-layersGPU offload layer count0Set to 99 to offload all layers to GPU
--threadsCPU threads for generationAutoSet explicitly for consistent performance
--threads-batchThreads for batch processingSame as --threadsTune separately for prompt vs generation
--cont-batchingContinuous batchingOffEnable for multiple concurrent requests
--parallelParallel sequence count1Increase for concurrent request handling
--mlockLock model in memoryOffPrevent swapping on systems with sufficient RAM
--embeddingEnable embeddings endpointOffRequired for /v1/embeddings API

Network-Accessible Configuration

To allow connections from other machines (development/testing only):

./llamafile --server \
    -m /path/to/model.gguf \
    --nobrowser \
    --host 0.0.0.0 \
    --port 8080

Security warning: Binding to 0.0.0.0 exposes the API to network access. Use only in trusted environments.

API Integration

Using LiteLLM (Recommended)

LiteLLM provides unified interface for llamafile and cloud LLM providers.

import litellm

response = litellm.completion(
    model="llamafile/gemma-3-3b",  # MUST use llamafile/ prefix
    messages=[{"role": "user", "content": "Hello, world!"}],
    api_base="http://localhost:8080/v1",  # MUST include /v1 suffix
    temperature=0.3,
    max_tokens=200,
)

print(response.choices[0].message.content)

Critical requirements for LiteLLM:

  1. Model name MUST use llamafile/ prefix for routing
  2. api_base MUST include /v1 suffix
  3. No API key required (any placeholder value works)

Related skill: For comprehensive LiteLLM configuration, activate the litellm skill:

Skill(skill: "litellm:litellm")

Using OpenAI Python SDK

Direct integration with OpenAI SDK for llamafile endpoints:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8080/v1",  # MUST include /v1
    api_key="sk-no-key-required",  # Any value works
)

response = client.chat.completions.create(
    model="local-model",  # Model name is flexible
    messages=[{"role": "user", "content": "Hello, world!"}],
    temperature=0.3,
    max_tokens=200,
)

print(response.choices[0].message.content)

Using curl for Testing

Verify llamafile server is responding correctly:

# Health check
curl http://localhost:8080/health

# Chat completions
curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "local",
    "messages": [{"role": "user", "content": "Hello"}],
    "temperature": 0.3,
    "max_tokens": 200
  }'

# Embeddings (requires --embedding flag on server)
curl http://localhost:8080/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{
    "model": "local",
    "input": ["Hello world"]
  }'

Server Management

Process Management Script

Python script to start llamafile as background process with health checking:

import subprocess
import time
import httpx


def start_llamafile(
    llamafile_path: str, model_path: str, port: int = 8080, host: str = "127.0.0.1"
) -> subprocess.Popen:
    """Start llamafile server as background process."""
    cmd = [llamafile_path, "--server", "-m", model_path, "--nobrowser", "--port", str(port), "--host", host]
    process = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)
    _wait_for_server(host, port)
    return process


def _wait_for_server(host: str, port: int, timeout: int = 30) -> None:
    """Wait for server to respond to health checks."""
    url = f"http://{host}:{port}/health"
    start = time.time()
    while time.time() - start < timeout:
        try:
            response = httpx.get(url, timeout=2)
            if response.status_code == 200:
                return
        except httpx.RequestError:
            pass
        time.sleep(0.5)
    raise TimeoutError(f"Server did not start within {timeout} seconds")

Configuration File Pattern

Example TOML configuration for applications using llamafile:

# ~/.config/app-name/config.toml
[ai]
model = "llamafile/gemma-3-3b"  # Must use llamafile/ prefix
temperature = 0.3
max_tokens = 200

[llamafile]
path = "/home/user/.local/bin/llamafile"
model_path = "/home/user/.local/share/app-name/models/gemma-3-3b.gguf"
api_base = "http://127.0.0.1:8080/v1"  # Include /v1 suffix

Troubleshooting

Server Fails to Start

Check if port is already in use:

# Find process using port 8080
lsof -i :8080

# Kill existing process
kill $(lsof -t -i :8080)

Verify model file exists and is readable:

ls -lh /path/to/model.gguf

Check llamafile binary permissions:

ls -la /path/to/llamafile
# Should show: -rwxr-xr-x (executable)

# Fix permissions if needed
chmod 755 /path/to/llamafile

Connection Refused Errors

Verify server is running:

# Check health endpoint
curl http://localhost:8080/health

# Check server is listening
netstat -tlnp | grep 8080
# or
lsof -i :8080

Common causes:

  1. Server not started with --server flag
  2. Wrong port number (8080 vs 8000)
  3. Missing /v1 in API URL path
  4. Server bound to 127.0.0.1 but accessing from another machine

API Errors

Test basic connectivity:

# Verbose health check
curl -v http://localhost:8080/health

# Test chat completions with verbose output
curl -v http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"test","messages":[{"role":"user","content":"Hi"}]}'

Common API issues:

ErrorCauseSolution
404 Not FoundMissing /v1 in URLAdd /v1 before endpoint path
Connection refusedServer not runningStart server with --server flag
TimeoutModel loading slowlyWait longer or use smaller model
Invalid modelWrong model pathVerify -m path to GGUF file

Performance Issues

Optimize inference speed:

  1. Use quantized models (Q4_K_M recommended)
  2. Enable GPU acceleration: --n-gpu-layers 99
  3. Increase threads: --threads 8
  4. Enable continuous batching: --cont-batching
  5. Reduce context size if not needed: --ctx-size 2048

Check GPU availability:

# NVIDIA GPU
nvidia-smi

# AMD GPU
rocm-smi

# Apple Metal (check activity monitor)

Common Pitfalls

Avoid these frequent errors when using llamafile:

  1. Port 8000 vs 8080: Llamafile defaults to port 8080, not 8000
  2. Missing /v1 in API URL: Always include /v1 suffix for OpenAI-compatible endpoints
  3. LiteLLM prefix: Must use llamafile/ prefix in model name for proper routing
  4. API key confusion: No real API key needed, but some clients require placeholder value
  5. Starting server from hooks: Application hooks should check if server is running, not start it
  6. Model path issues: Ensure GGUF file exists and is readable before starting server
  7. Binary permissions: Llamafile must be executable (chmod 755)
  8. GPU layers on CPU: Setting --n-gpu-layers on CPU-only systems causes errors

Version Information

Current stable version: 0.9.3 (May 14, 2025)

Version constants:

LLAMAFILE_MAJOR = 0
LLAMAFILE_MINOR = 9
LLAMAFILE_PATCH = 3

Recent changes in 0.9.3:

  • Added Phi4 model support
  • Added Qwen3 model support
  • Respects NO_COLOR environment variable
  • Fixed URL handling in JavaScript (preserves path when building relative URLs)
  • Added Plaintext output option to LocalScore

Related Skills and Tools

Skills to activate:

  • litellm - For unified LLM provider interface and routing
    Skill(skill: "litellm:litellm")
    

External tools:

  • LiteLLM - Unified interface for multiple LLM providers
  • OpenAI Python SDK - Direct OpenAI-compatible API access
  • llama.cpp - Underlying inference engine
  • GGUF format - Model format specification

References

Official Documentation

Model Resources

Related Technologies

Frequently asked questions

What to verify before installation and use

What does the llamafile source document cover?

Configure and manage Mozilla Llamafile - a cross-platform executable distribution format that runs LLMs locally with an OpenAI-compatible API.

How do I install llamafile?

The source record exposes this install command: npx skills add https://github.com/Jamie-BitFlight/claude_skills --skill "plugins/llamafile/skills/llamafile". Inspect the command and pinned source before running it.

Which permission-related actions were detected?

Static rules flagged write-files, network, exec-script in the source; the page lists the matching lines and excerpts.

Alternatives

Compare before choosing

Computed 10045,960

coreyhaines31/marketingskills

ab-testing

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program

Computed 10029,236

garrytan/gbrain

bulk-ingestion

End-to-end discipline for turning any large data source (audio libraries, email takeouts, document corpora, chat exports, API dumps) into brain pages at scale. The lifecycle spine: SCHEMA → ACCESS → TRIAL → EVALUATE → IMPROVE → CODIFY → TEST → SKILLIFY → BULK → MONITOR. State is tracked in a durable JSON manifest (see MANIFEST-PATTERN.md) so any crash, session boundary, or subagent fan-out resumes from ground truth instead of memory.

Computed 10025,136

alirezarezvani/claude-skills

app-store-optimization

App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store. Use when the user asks about ASO, app store rankings, app metadata, app titles and descriptions, app store listings, app visibility, or mobile app marketing on iOS or Android. Supports keyword research and scoring, competitor keyword analysis, metadata optimization, A/B test planning, launch checklist

Computed 1005,277

dotnet/skills

migrate-vstest-to-mtp

Migrates .NET test projects from VSTest to Microsoft.Testing.Platform (MTP). Use when user asks to "migrate to MTP", "switch from VSTest", "enable Microsoft.Testing.Platform", "use MTP runner", set OutputType=Exe only for test projects in Directory.Build.props, or mentions EnableMSTestRunner, EnableNUnitRunner, or UseMicrosoftTestingPlatformRunner. USE FOR: MTP behavioral differences vs VSTest (exit code 8, zero tests discovered, --ignore-exit-code, TESTINGPLATFORM_EXITCODE_IGNORE); centralizing