Ollama Local LLM Skill

SkillAI & models

Run local LLM inference for chat, text generation, and embeddings via the Ollama server at {{OLLAMA_HOST}}:{{OLLAMA_PORT}}.

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Ollama Local LLM Skill skill

What this skill tells your AI

The instructions your AI receives, as published by bidewio/better-openclaw in skills/ollama-local-llm/SKILL.md and read by ahel’s review.

Ollama local LLM server is available at http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}} within the Docker network.

Chat Completion

Send a multi-turn conversation and get a response:

curl -X POST "http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/api/chat" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.2",
    "messages": [
      { "role": "system", "content": "You are a helpful coding assistant." },
      { "role": "user", "content": "Write a Python function to reverse a string." }
    ],
    "stream": false
  }'

Text Generation

Generate text from a single prompt:

curl -X POST "http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/api/generate" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.2",
    "prompt": "Explain the concept of recursion in simple terms.",
    "stream": false
  }'

Streaming Responses

For real-time token-by-token output, enable streaming:

curl -X POST "http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/api/chat" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.2",
    "messages": [
      { "role": "user", "content": "Tell me a short story." }
    ],
    "stream": true
  }'

Generating Embeddings

Create vector embeddings for text (useful with Qdrant):

curl -X POST "http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/api/embed" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nomic-embed-text",
    "input": ["This is a sentence to embed.", "Another sentence for comparison."]
  }'

Model Management

# List available models
curl "http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/api/tags"

# Pull a new model
curl -X POST "http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/api/pull" \
  -H "Content-Type: application/json" \
  -d '{"name": "llama3.2"}'

# Show model details
curl -X POST "http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/api/show" \
  -H "Content-Type: application/json" \
  -d '{"name": "llama3.2"}'

# Delete a model
curl -X DELETE "http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/api/delete" \
  -H "Content-Type: application/json" \
  -d '{"name": "old-model"}'

Advanced Generation Options

Fine-tune generation with parameters:

curl -X POST "http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/api/chat" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.2",
    "messages": [
      { "role": "user", "content": "Write a creative poem about the ocean." }
    ],
    "stream": false,
    "options": {
      "temperature": 0.8,
      "top_p": 0.9,
      "top_k": 40,
      "num_predict": 512,
      "seed": 42
    }
  }'

Using a Custom System Prompt

curl -X POST "http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/api/chat" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.2",
    "messages": [
      { "role": "system", "content": "You are an expert data analyst. Respond with structured JSON when possible." },
      { "role": "user", "content": "Analyze this data: [10, 25, 18, 42, 7, 33]" }
    ],
    "stream": false,
    "format": "json"
  }'

Recommended Models

ModelUse CaseSize
llama3.2General chat and reasoning3B
llama3.2:70bComplex reasoning tasks70B
codellamaCode generation and review7B
nomic-embed-textText embeddings for RAG137M
mistralFast general-purpose inference7B
phi3Compact and efficient reasoning3.8B

Tips for AI Agents

  • Always set "stream": false when you need to parse the complete response programmatically.
  • Use format: "json" when you need structured output that's easy to parse.
  • Check available models with /api/tags before making inference calls to avoid 404 errors.
  • For embedding tasks, use nomic-embed-text or similar dedicated embedding models, not chat models.
  • Lower temperature (0.1-0.3) for factual/deterministic tasks; raise it (0.7-1.0) for creative tasks.
  • Set num_predict to limit response length and prevent runaway generation.
  • Use the embeddings endpoint with Qdrant for building RAG (Retrieval-Augmented Generation) pipelines.
  • Check Ollama health at http://{{OLLAMA_HOST}}:{{OLLAMA_PORT}}/ — it returns "Ollama is running".

Signals

GitHub stars
58
Forks
6
Last commit
Aug 2026
Advanced
Item type
skill
Key
ollama-local-llm
Source
github.com/bidewio/better-openclaw