Initial commit: arr-stack homelab + local AI
Docker Compose media stack (gluetun, *arr, Jellyfin), Ternary-Bonsai AI (Open WebUI, Forgejo, optional OpenHands), and Ansible bootstrap for the SER5 Ubuntu host.
This commit is contained in:
commit
3645d1314c
33 changed files with 2673 additions and 0 deletions
323
README.md
Normal file
323
README.md
Normal file
|
|
@ -0,0 +1,323 @@
|
|||
# Arr Stack + Local AI
|
||||
|
||||
A dockerized media automation stack (VPN-locked torrents + Jellyfin) **plus** a
|
||||
local AI stack built around
|
||||
[Ternary-Bonsai-27B](https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf)
|
||||
for the **Beelink SER5 PRO** (Ryzen 7 7735HS + Radeon 680M).
|
||||
|
||||
## Services
|
||||
|
||||
| Service | URL | Purpose |
|
||||
|-------------|------------------------|--------------------------------------|
|
||||
| Jellyfin | http://host:8096 | Media server / playback |
|
||||
| qBittorrent | http://host:8080 | Torrent client (VPN-only) |
|
||||
| Prowlarr | http://host:9696 | Indexer manager |
|
||||
| Sonarr | http://host:8989 | TV shows |
|
||||
| Radarr | http://host:7878 | Movies |
|
||||
| Bazarr | http://host:6767 | Subtitles |
|
||||
| Lidarr | http://host:8686 | Music |
|
||||
| **Forgejo** | http://host:3002 | Self-hosted Git (SSH :2222) |
|
||||
| **Open WebUI** | http://host:3000 | Chat UI, RAG, tools, code interpreter |
|
||||
| **Bonsai API** | http://host:8081/v1 | OpenAI-compatible llama-server |
|
||||
| SearXNG | http://127.0.0.1:8888 | Private web search (for RAG) |
|
||||
| OpenHands | http://host:3001 | Coding agent (`--profile coding`) |
|
||||
|
||||
---
|
||||
|
||||
## Media stack
|
||||
|
||||
### 1. Prepare
|
||||
|
||||
```bash
|
||||
cd arr-stack
|
||||
cp .env.example .env # then edit .env (PUID/PGID, TZ, LAN_SUBNET)
|
||||
```
|
||||
|
||||
Drop your provider's WireGuard config into `wireguard/wg0.conf` and make the
|
||||
two edits described in `wireguard/wg0.conf.example`:
|
||||
|
||||
1. **Remove IPv6 addresses** from `Address` / `DNS` lines.
|
||||
2. **Endpoint must be an IP**, not a hostname (`dig +short your.endpoint.host`).
|
||||
|
||||
Create the data tree (already scaffolded here; on a fresh host run):
|
||||
|
||||
```bash
|
||||
mkdir -p config data/media/{movies,tv,music} data/torrents/{movies,tv,music}
|
||||
```
|
||||
|
||||
The layout uses a **single `/data` share** — this lets the Arrs hardlink
|
||||
finished downloads into the media library instantly (no copy, no double disk
|
||||
usage, torrent keeps seeding):
|
||||
|
||||
```
|
||||
data/
|
||||
├── media/ <- Jellyfin libraries
|
||||
│ ├── movies/
|
||||
│ ├── tv/
|
||||
│ └── music/
|
||||
└── torrents/ <- qBittorrent downloads here
|
||||
├── movies/
|
||||
├── tv/
|
||||
└── music/
|
||||
```
|
||||
|
||||
### 2. Start media
|
||||
|
||||
```bash
|
||||
docker compose up -d
|
||||
docker compose ps # gluetun must show (healthy)
|
||||
docker logs -f gluetun # wait for "You are running on the VPN" style line
|
||||
```
|
||||
|
||||
qBittorrent waits for gluetun to be healthy before starting.
|
||||
|
||||
### 3. Verify the lockdown (do this once)
|
||||
|
||||
```bash
|
||||
# IP seen from inside qBittorrent — must be your VPN's IP, NOT your home IP:
|
||||
docker exec qbittorrent curl -s ifconfig.me && echo
|
||||
|
||||
# Your real IP for comparison:
|
||||
curl -s ifconfig.me && echo
|
||||
|
||||
# Killswitch test — kill the tunnel, traffic must die with it:
|
||||
docker exec gluetun sh -c "ip link del tun0" 2>/dev/null
|
||||
docker exec qbittorrent curl -s --max-time 8 ifconfig.me || echo "NO CONNECTIVITY (good)"
|
||||
docker restart gluetun # bring the tunnel back
|
||||
```
|
||||
|
||||
Belt-and-suspenders extra: in qBittorrent → Tools → Options → Advanced, set
|
||||
**Network interface** to `tun0`. Even in a hypothetical gluetun failure,
|
||||
qBittorrent itself would then refuse to use any other interface.
|
||||
|
||||
### 4. Wire it together
|
||||
|
||||
**qBittorrent** (http://host:8080 — default user `admin`, temporary password
|
||||
is printed in `docker logs qbittorrent` on first start; change it):
|
||||
|
||||
- Options → Downloads → Default Save Path: `/data/torrents`
|
||||
- Create categories `movies` → `/data/torrents/movies` and `tv` → `/data/torrents/tv`
|
||||
|
||||
**Prowlarr** (http://host:9696):
|
||||
|
||||
- Add your indexers under *Indexers*.
|
||||
- *Settings → Apps*: add Sonarr (`http://sonarr:8989`), Radarr
|
||||
(`http://radarr:7878`), each with its API key (found in the Arr under
|
||||
*Settings → General*). Indexers sync automatically from then on.
|
||||
|
||||
**Sonarr & Radarr** — *Settings → Download Clients* → add qBittorrent:
|
||||
|
||||
- Host: `gluetun`, Port: `8080`, your qBittorrent credentials
|
||||
- Category: `tv` (Sonarr) / `movies` (Radarr)
|
||||
- *Settings → Media Management*: root folder `/data/media/tv` or
|
||||
`/data/media/movies`; **enable hardlinks** (on by default with this layout)
|
||||
|
||||
**Bazarr**: point it at Sonarr/Radarr (`http://sonarr:8989` /
|
||||
`http://radarr:7878` + API keys); paths `/data/media/tv`, `/data/media/movies`.
|
||||
|
||||
**Lidarr** (http://host:8686) — music, same patterns as Sonarr/Radarr:
|
||||
|
||||
- *Settings → Download Clients* → qBittorrent: Host `gluetun`, Port `8080`
|
||||
- Category: `music` → qBittorrent save path `/data/torrents/music`
|
||||
- *Settings → Media Management*: root folder `/data/media/music`; hardlinks on
|
||||
- In **Prowlarr** → *Settings → Apps*: add Lidarr (`http://lidarr:8686` + API key)
|
||||
|
||||
**Jellyfin** (http://host:8096): add libraries → Movies: `/data/media/movies`,
|
||||
Shows: `/data/media/tv`, Music: `/data/media/music`. Its media mount is
|
||||
read-only; enable trickplay/metadata into `/config` (the default) — it never
|
||||
writes to your media.
|
||||
|
||||
### 5. Optional: provider port forwarding
|
||||
|
||||
If your VPN provider gives you a forwarded port (AirVPN, ProtonVPN, …),
|
||||
inbound peers can reach you and swarm connectivity improves:
|
||||
|
||||
1. Uncomment `FIREWALL_VPN_INPUT_PORTS` in the compose file and set the port.
|
||||
2. For providers with a NAT-PMP API also uncomment `VPN_PORT_FORWARDING` and
|
||||
`VPN_PORT_FORWARDING_PROVIDER`.
|
||||
3. Set the same port as qBittorrent's listening port (Options → Connection)
|
||||
and disable UPnP/NAT-PMP inside qBittorrent.
|
||||
|
||||
Never publish qBittorrent's BitTorrent port on the host — inbound connections
|
||||
must arrive through the tunnel only.
|
||||
|
||||
---
|
||||
|
||||
## Local AI stack
|
||||
|
||||
### Why this design (and not plain Ollama)
|
||||
|
||||
| Piece | Choice | Reason |
|
||||
|-------|--------|--------|
|
||||
| Model | **Ternary-Bonsai-27B** (~7.2 GB Q2_0) | ~95% of FP16 quality at laptop footprint; strong coding/math |
|
||||
| Runtime | **PrismML llama.cpp** (`llama-server`) | Custom Q2_0 g128 kernels — stock Ollama / mainline won't load the default GGUF |
|
||||
| Hardware | **SER5 PRO** Ryzen 7 7735HS + Radeon 680M | No NVIDIA → default is **CPU**; optional Vulkan for the iGPU |
|
||||
| Web UI | **Open WebUI** | Chat, document RAG, web search, tools, code interpreter |
|
||||
| Search | **SearXNG** | Keyless private search for RAG |
|
||||
| Coding | **OpenHands** (optional profile) + `/workspace` mount | Agent that edits files / runs commands in a sandbox |
|
||||
|
||||
**Memory ballpark** (model alone, from Prism docs): ~8 GB at 4–10k context, ~14 GB at 100k with FP16 KV. With Arr + Jellyfin running, **32 GB system RAM is the practical minimum**; **64 GB** is comfortable for long context + media.
|
||||
|
||||
### 1. Download the model (~7.2 GB)
|
||||
|
||||
```bash
|
||||
# Recommended (resume-friendly):
|
||||
pip install -U "huggingface_hub[cli]"
|
||||
|
||||
./scripts/download-bonsai-model.sh
|
||||
# files land in ./models/bonsai/
|
||||
```
|
||||
|
||||
Optional:
|
||||
|
||||
```bash
|
||||
DOWNLOAD_DRAFTER=1 ./scripts/download-bonsai-model.sh # speculative decoding pack (CUDA-oriented)
|
||||
```
|
||||
|
||||
### 2. Start AI services
|
||||
|
||||
```bash
|
||||
# Build Prism llama-server image + start UI + search
|
||||
docker compose up -d --build bonsai open-webui searxng
|
||||
|
||||
# Health
|
||||
curl -s http://localhost:8081/health
|
||||
curl -s http://localhost:8081/v1/models
|
||||
```
|
||||
|
||||
Open **http://host:3000** — create the admin account on first visit.
|
||||
Open WebUI is pre-pointed at `http://bonsai:8080/v1` (see compose env).
|
||||
|
||||
### 3. RAG (documents & local files)
|
||||
|
||||
**Upload / Knowledge collections (best for docs)**
|
||||
|
||||
1. Open WebUI → **Workspace → Knowledge** → create a collection.
|
||||
2. Upload PDFs, markdown, code, notes.
|
||||
3. In a chat, attach the collection (or enable it on the model).
|
||||
|
||||
Embeddings download into the Open WebUI volume on first use (local, no cloud).
|
||||
|
||||
**Whole folders via bind-mount**
|
||||
|
||||
`AI_WORKSPACE_DIR` (default `./workspace`) is mounted **read-only** at
|
||||
`/workspace` inside Open WebUI. Put projects/notes there, or point
|
||||
`.env` at your real code tree:
|
||||
|
||||
```bash
|
||||
# .env
|
||||
AI_WORKSPACE_DIR=/home/you/src
|
||||
```
|
||||
|
||||
Then ask e.g. “Summarize the layout of `/workspace/my-app`” after adding those
|
||||
paths to a Knowledge collection, or use tools/OpenHands for active editing.
|
||||
|
||||
**Web RAG**
|
||||
|
||||
SearXNG is wired as the search engine (`ENABLE_RAG_WEB_SEARCH=true`). In chat
|
||||
settings enable web search when you want live results.
|
||||
|
||||
### 4. Coding helpers
|
||||
|
||||
| Tool | How | Best for |
|
||||
|------|-----|----------|
|
||||
| **Open WebUI + Knowledge** | Upload repo docs / mount `/workspace` | Q&A over code, design docs |
|
||||
| **Open WebUI code interpreter** | Enabled in compose | Python plots, quick analysis |
|
||||
| **OpenHands** | `docker compose --profile coding up -d openhands` | Multi-step “change the code / run tests” agents |
|
||||
| **Continue / Aider / Cursor** (host apps) | Base URL `http://host:8081/v1`, any API key | IDE-native completion & edits |
|
||||
|
||||
**OpenHands** (http://host:3001):
|
||||
|
||||
1. Settings → enable Advanced.
|
||||
2. Custom model: `openai/<id-from-/v1/models>` (or keep the pre-set env).
|
||||
3. Base URL: `http://bonsai:8080/v1` (from inside Docker) or `http://host.docker.internal:8081/v1`.
|
||||
4. API key: same as `OPENAI_API_KEY` in `.env`.
|
||||
5. Workspace is `${AI_WORKSPACE_DIR}` — agent sandboxes use Docker socket.
|
||||
|
||||
> Note: Prism lists long-horizon multi-file agentic coding as a weaker area of
|
||||
> this 27B release; it is still very usable for review, refactors, and smaller
|
||||
> tasks. Prefer shorter goals and verify with tests.
|
||||
|
||||
### 5. Performance knobs (SER5 PRO)
|
||||
|
||||
```bash
|
||||
# .env examples
|
||||
|
||||
# Leave headroom for Jellyfin/Arrs on a 16-thread 7735HS
|
||||
BONSAI_THREADS=12
|
||||
|
||||
# Longer context for big repos / PDFs (needs RAM)
|
||||
BONSAI_CTX=32768
|
||||
BONSAI_KV4=1 # 4-bit KV cache if you approach OOM
|
||||
|
||||
# Optional: Vulkan offload on Radeon 680M
|
||||
BONSAI_BACKEND=vulkan
|
||||
BONSAI_NGL=99
|
||||
# then uncomment devices/group_add under bonsai in docker-compose.yml
|
||||
docker compose build --no-cache bonsai && docker compose up -d bonsai
|
||||
```
|
||||
|
||||
Start with **CPU (`BONSAI_NGL=0`)**. iGPU Vulkan can help but shares system
|
||||
RAM and is more fiddly in Docker; measure both with a short prompt if you care.
|
||||
|
||||
Recommended generation defaults (already set): `temp=0.7`, `top_p=0.95`,
|
||||
`top_k=20`. The 27B is a **thinking** model — answers may include a reasoning
|
||||
block; that is expected and improves quality.
|
||||
|
||||
### 6. Quick API smoke test
|
||||
|
||||
```bash
|
||||
curl -s http://localhost:8081/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer sk-local-bonsai" \
|
||||
-d '{
|
||||
"model": "local",
|
||||
"messages": [{"role":"user","content":"Say hello in one short sentence."}],
|
||||
"temperature": 0.7,
|
||||
"max_tokens": 128
|
||||
}'
|
||||
```
|
||||
|
||||
(`model` can be any string or the id returned by `/v1/models` — llama-server is
|
||||
usually lenient.)
|
||||
|
||||
### 7. What not to do
|
||||
|
||||
- Do **not** route bonsai/open-webui through gluetun — inference is local and
|
||||
the VPN killswitch would only add pain.
|
||||
- Do **not** expect stock Ollama to load `Ternary-Bonsai-27B-Q2_0.gguf` without
|
||||
Prism’s kernels (use our `bonsai` service).
|
||||
- Keep Open WebUI / OpenHands off the public internet; use LAN + auth (`WEBUI_AUTH=true`).
|
||||
|
||||
---
|
||||
|
||||
## Layout after AI setup
|
||||
|
||||
```
|
||||
arr-stack/
|
||||
├── docker-compose.yml
|
||||
├── .env / .env.example
|
||||
├── ai/ # Prism llama-server image
|
||||
│ ├── Dockerfile
|
||||
│ └── entrypoint.sh
|
||||
├── models/bonsai/ # GGUF weights (gitignored content)
|
||||
├── workspace/ # files for RAG / OpenHands
|
||||
├── searxng/settings.yml
|
||||
├── scripts/download-bonsai-model.sh
|
||||
├── config/ # app configs (open-webui, sonarr, …)
|
||||
├── data/ # media + torrents
|
||||
└── wireguard/wg0.conf
|
||||
```
|
||||
|
||||
## Notes
|
||||
|
||||
- All app-to-app traffic uses container DNS names (`gluetun`, `sonarr`,
|
||||
`bonsai`, …) on the internal `arr-net` bridge; only WebUIs are exposed to
|
||||
your LAN.
|
||||
- Update media stack with `docker compose pull && docker compose up -d`.
|
||||
Gluetun is pinned to major version `v3` — read its release notes before
|
||||
jumping majors.
|
||||
- Rebuild AI runtime after Prism release bumps: edit `PRISM_RELEASE` in
|
||||
`ai/Dockerfile`, then `docker compose build --no-cache bonsai`.
|
||||
- Optional media extras (commented in compose): Jellyseerr.
|
||||
- Optional coding profile: `docker compose --profile coding up -d openhands`.
|
||||
Loading…
Add table
Add a link
Reference in a new issue