324 lines
12 KiB
Markdown
324 lines
12 KiB
Markdown
|
|
# Arr Stack + Local AI
|
|||
|
|
|
|||
|
|
A dockerized media automation stack (VPN-locked torrents + Jellyfin) **plus** a
|
|||
|
|
local AI stack built around
|
|||
|
|
[Ternary-Bonsai-27B](https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf)
|
|||
|
|
for the **Beelink SER5 PRO** (Ryzen 7 7735HS + Radeon 680M).
|
|||
|
|
|
|||
|
|
## Services
|
|||
|
|
|
|||
|
|
| Service | URL | Purpose |
|
|||
|
|
|-------------|------------------------|--------------------------------------|
|
|||
|
|
| Jellyfin | http://host:8096 | Media server / playback |
|
|||
|
|
| qBittorrent | http://host:8080 | Torrent client (VPN-only) |
|
|||
|
|
| Prowlarr | http://host:9696 | Indexer manager |
|
|||
|
|
| Sonarr | http://host:8989 | TV shows |
|
|||
|
|
| Radarr | http://host:7878 | Movies |
|
|||
|
|
| Bazarr | http://host:6767 | Subtitles |
|
|||
|
|
| Lidarr | http://host:8686 | Music |
|
|||
|
|
| **Forgejo** | http://host:3002 | Self-hosted Git (SSH :2222) |
|
|||
|
|
| **Open WebUI** | http://host:3000 | Chat UI, RAG, tools, code interpreter |
|
|||
|
|
| **Bonsai API** | http://host:8081/v1 | OpenAI-compatible llama-server |
|
|||
|
|
| SearXNG | http://127.0.0.1:8888 | Private web search (for RAG) |
|
|||
|
|
| OpenHands | http://host:3001 | Coding agent (`--profile coding`) |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Media stack
|
|||
|
|
|
|||
|
|
### 1. Prepare
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
cd arr-stack
|
|||
|
|
cp .env.example .env # then edit .env (PUID/PGID, TZ, LAN_SUBNET)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Drop your provider's WireGuard config into `wireguard/wg0.conf` and make the
|
|||
|
|
two edits described in `wireguard/wg0.conf.example`:
|
|||
|
|
|
|||
|
|
1. **Remove IPv6 addresses** from `Address` / `DNS` lines.
|
|||
|
|
2. **Endpoint must be an IP**, not a hostname (`dig +short your.endpoint.host`).
|
|||
|
|
|
|||
|
|
Create the data tree (already scaffolded here; on a fresh host run):
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
mkdir -p config data/media/{movies,tv,music} data/torrents/{movies,tv,music}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The layout uses a **single `/data` share** — this lets the Arrs hardlink
|
|||
|
|
finished downloads into the media library instantly (no copy, no double disk
|
|||
|
|
usage, torrent keeps seeding):
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
data/
|
|||
|
|
├── media/ <- Jellyfin libraries
|
|||
|
|
│ ├── movies/
|
|||
|
|
│ ├── tv/
|
|||
|
|
│ └── music/
|
|||
|
|
└── torrents/ <- qBittorrent downloads here
|
|||
|
|
├── movies/
|
|||
|
|
├── tv/
|
|||
|
|
└── music/
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 2. Start media
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
docker compose up -d
|
|||
|
|
docker compose ps # gluetun must show (healthy)
|
|||
|
|
docker logs -f gluetun # wait for "You are running on the VPN" style line
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
qBittorrent waits for gluetun to be healthy before starting.
|
|||
|
|
|
|||
|
|
### 3. Verify the lockdown (do this once)
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
# IP seen from inside qBittorrent — must be your VPN's IP, NOT your home IP:
|
|||
|
|
docker exec qbittorrent curl -s ifconfig.me && echo
|
|||
|
|
|
|||
|
|
# Your real IP for comparison:
|
|||
|
|
curl -s ifconfig.me && echo
|
|||
|
|
|
|||
|
|
# Killswitch test — kill the tunnel, traffic must die with it:
|
|||
|
|
docker exec gluetun sh -c "ip link del tun0" 2>/dev/null
|
|||
|
|
docker exec qbittorrent curl -s --max-time 8 ifconfig.me || echo "NO CONNECTIVITY (good)"
|
|||
|
|
docker restart gluetun # bring the tunnel back
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Belt-and-suspenders extra: in qBittorrent → Tools → Options → Advanced, set
|
|||
|
|
**Network interface** to `tun0`. Even in a hypothetical gluetun failure,
|
|||
|
|
qBittorrent itself would then refuse to use any other interface.
|
|||
|
|
|
|||
|
|
### 4. Wire it together
|
|||
|
|
|
|||
|
|
**qBittorrent** (http://host:8080 — default user `admin`, temporary password
|
|||
|
|
is printed in `docker logs qbittorrent` on first start; change it):
|
|||
|
|
|
|||
|
|
- Options → Downloads → Default Save Path: `/data/torrents`
|
|||
|
|
- Create categories `movies` → `/data/torrents/movies` and `tv` → `/data/torrents/tv`
|
|||
|
|
|
|||
|
|
**Prowlarr** (http://host:9696):
|
|||
|
|
|
|||
|
|
- Add your indexers under *Indexers*.
|
|||
|
|
- *Settings → Apps*: add Sonarr (`http://sonarr:8989`), Radarr
|
|||
|
|
(`http://radarr:7878`), each with its API key (found in the Arr under
|
|||
|
|
*Settings → General*). Indexers sync automatically from then on.
|
|||
|
|
|
|||
|
|
**Sonarr & Radarr** — *Settings → Download Clients* → add qBittorrent:
|
|||
|
|
|
|||
|
|
- Host: `gluetun`, Port: `8080`, your qBittorrent credentials
|
|||
|
|
- Category: `tv` (Sonarr) / `movies` (Radarr)
|
|||
|
|
- *Settings → Media Management*: root folder `/data/media/tv` or
|
|||
|
|
`/data/media/movies`; **enable hardlinks** (on by default with this layout)
|
|||
|
|
|
|||
|
|
**Bazarr**: point it at Sonarr/Radarr (`http://sonarr:8989` /
|
|||
|
|
`http://radarr:7878` + API keys); paths `/data/media/tv`, `/data/media/movies`.
|
|||
|
|
|
|||
|
|
**Lidarr** (http://host:8686) — music, same patterns as Sonarr/Radarr:
|
|||
|
|
|
|||
|
|
- *Settings → Download Clients* → qBittorrent: Host `gluetun`, Port `8080`
|
|||
|
|
- Category: `music` → qBittorrent save path `/data/torrents/music`
|
|||
|
|
- *Settings → Media Management*: root folder `/data/media/music`; hardlinks on
|
|||
|
|
- In **Prowlarr** → *Settings → Apps*: add Lidarr (`http://lidarr:8686` + API key)
|
|||
|
|
|
|||
|
|
**Jellyfin** (http://host:8096): add libraries → Movies: `/data/media/movies`,
|
|||
|
|
Shows: `/data/media/tv`, Music: `/data/media/music`. Its media mount is
|
|||
|
|
read-only; enable trickplay/metadata into `/config` (the default) — it never
|
|||
|
|
writes to your media.
|
|||
|
|
|
|||
|
|
### 5. Optional: provider port forwarding
|
|||
|
|
|
|||
|
|
If your VPN provider gives you a forwarded port (AirVPN, ProtonVPN, …),
|
|||
|
|
inbound peers can reach you and swarm connectivity improves:
|
|||
|
|
|
|||
|
|
1. Uncomment `FIREWALL_VPN_INPUT_PORTS` in the compose file and set the port.
|
|||
|
|
2. For providers with a NAT-PMP API also uncomment `VPN_PORT_FORWARDING` and
|
|||
|
|
`VPN_PORT_FORWARDING_PROVIDER`.
|
|||
|
|
3. Set the same port as qBittorrent's listening port (Options → Connection)
|
|||
|
|
and disable UPnP/NAT-PMP inside qBittorrent.
|
|||
|
|
|
|||
|
|
Never publish qBittorrent's BitTorrent port on the host — inbound connections
|
|||
|
|
must arrive through the tunnel only.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Local AI stack
|
|||
|
|
|
|||
|
|
### Why this design (and not plain Ollama)
|
|||
|
|
|
|||
|
|
| Piece | Choice | Reason |
|
|||
|
|
|-------|--------|--------|
|
|||
|
|
| Model | **Ternary-Bonsai-27B** (~7.2 GB Q2_0) | ~95% of FP16 quality at laptop footprint; strong coding/math |
|
|||
|
|
| Runtime | **PrismML llama.cpp** (`llama-server`) | Custom Q2_0 g128 kernels — stock Ollama / mainline won't load the default GGUF |
|
|||
|
|
| Hardware | **SER5 PRO** Ryzen 7 7735HS + Radeon 680M | No NVIDIA → default is **CPU**; optional Vulkan for the iGPU |
|
|||
|
|
| Web UI | **Open WebUI** | Chat, document RAG, web search, tools, code interpreter |
|
|||
|
|
| Search | **SearXNG** | Keyless private search for RAG |
|
|||
|
|
| Coding | **OpenHands** (optional profile) + `/workspace` mount | Agent that edits files / runs commands in a sandbox |
|
|||
|
|
|
|||
|
|
**Memory ballpark** (model alone, from Prism docs): ~8 GB at 4–10k context, ~14 GB at 100k with FP16 KV. With Arr + Jellyfin running, **32 GB system RAM is the practical minimum**; **64 GB** is comfortable for long context + media.
|
|||
|
|
|
|||
|
|
### 1. Download the model (~7.2 GB)
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
# Recommended (resume-friendly):
|
|||
|
|
pip install -U "huggingface_hub[cli]"
|
|||
|
|
|
|||
|
|
./scripts/download-bonsai-model.sh
|
|||
|
|
# files land in ./models/bonsai/
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Optional:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
DOWNLOAD_DRAFTER=1 ./scripts/download-bonsai-model.sh # speculative decoding pack (CUDA-oriented)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 2. Start AI services
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
# Build Prism llama-server image + start UI + search
|
|||
|
|
docker compose up -d --build bonsai open-webui searxng
|
|||
|
|
|
|||
|
|
# Health
|
|||
|
|
curl -s http://localhost:8081/health
|
|||
|
|
curl -s http://localhost:8081/v1/models
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Open **http://host:3000** — create the admin account on first visit.
|
|||
|
|
Open WebUI is pre-pointed at `http://bonsai:8080/v1` (see compose env).
|
|||
|
|
|
|||
|
|
### 3. RAG (documents & local files)
|
|||
|
|
|
|||
|
|
**Upload / Knowledge collections (best for docs)**
|
|||
|
|
|
|||
|
|
1. Open WebUI → **Workspace → Knowledge** → create a collection.
|
|||
|
|
2. Upload PDFs, markdown, code, notes.
|
|||
|
|
3. In a chat, attach the collection (or enable it on the model).
|
|||
|
|
|
|||
|
|
Embeddings download into the Open WebUI volume on first use (local, no cloud).
|
|||
|
|
|
|||
|
|
**Whole folders via bind-mount**
|
|||
|
|
|
|||
|
|
`AI_WORKSPACE_DIR` (default `./workspace`) is mounted **read-only** at
|
|||
|
|
`/workspace` inside Open WebUI. Put projects/notes there, or point
|
|||
|
|
`.env` at your real code tree:
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
# .env
|
|||
|
|
AI_WORKSPACE_DIR=/home/you/src
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Then ask e.g. “Summarize the layout of `/workspace/my-app`” after adding those
|
|||
|
|
paths to a Knowledge collection, or use tools/OpenHands for active editing.
|
|||
|
|
|
|||
|
|
**Web RAG**
|
|||
|
|
|
|||
|
|
SearXNG is wired as the search engine (`ENABLE_RAG_WEB_SEARCH=true`). In chat
|
|||
|
|
settings enable web search when you want live results.
|
|||
|
|
|
|||
|
|
### 4. Coding helpers
|
|||
|
|
|
|||
|
|
| Tool | How | Best for |
|
|||
|
|
|------|-----|----------|
|
|||
|
|
| **Open WebUI + Knowledge** | Upload repo docs / mount `/workspace` | Q&A over code, design docs |
|
|||
|
|
| **Open WebUI code interpreter** | Enabled in compose | Python plots, quick analysis |
|
|||
|
|
| **OpenHands** | `docker compose --profile coding up -d openhands` | Multi-step “change the code / run tests” agents |
|
|||
|
|
| **Continue / Aider / Cursor** (host apps) | Base URL `http://host:8081/v1`, any API key | IDE-native completion & edits |
|
|||
|
|
|
|||
|
|
**OpenHands** (http://host:3001):
|
|||
|
|
|
|||
|
|
1. Settings → enable Advanced.
|
|||
|
|
2. Custom model: `openai/<id-from-/v1/models>` (or keep the pre-set env).
|
|||
|
|
3. Base URL: `http://bonsai:8080/v1` (from inside Docker) or `http://host.docker.internal:8081/v1`.
|
|||
|
|
4. API key: same as `OPENAI_API_KEY` in `.env`.
|
|||
|
|
5. Workspace is `${AI_WORKSPACE_DIR}` — agent sandboxes use Docker socket.
|
|||
|
|
|
|||
|
|
> Note: Prism lists long-horizon multi-file agentic coding as a weaker area of
|
|||
|
|
> this 27B release; it is still very usable for review, refactors, and smaller
|
|||
|
|
> tasks. Prefer shorter goals and verify with tests.
|
|||
|
|
|
|||
|
|
### 5. Performance knobs (SER5 PRO)
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
# .env examples
|
|||
|
|
|
|||
|
|
# Leave headroom for Jellyfin/Arrs on a 16-thread 7735HS
|
|||
|
|
BONSAI_THREADS=12
|
|||
|
|
|
|||
|
|
# Longer context for big repos / PDFs (needs RAM)
|
|||
|
|
BONSAI_CTX=32768
|
|||
|
|
BONSAI_KV4=1 # 4-bit KV cache if you approach OOM
|
|||
|
|
|
|||
|
|
# Optional: Vulkan offload on Radeon 680M
|
|||
|
|
BONSAI_BACKEND=vulkan
|
|||
|
|
BONSAI_NGL=99
|
|||
|
|
# then uncomment devices/group_add under bonsai in docker-compose.yml
|
|||
|
|
docker compose build --no-cache bonsai && docker compose up -d bonsai
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Start with **CPU (`BONSAI_NGL=0`)**. iGPU Vulkan can help but shares system
|
|||
|
|
RAM and is more fiddly in Docker; measure both with a short prompt if you care.
|
|||
|
|
|
|||
|
|
Recommended generation defaults (already set): `temp=0.7`, `top_p=0.95`,
|
|||
|
|
`top_k=20`. The 27B is a **thinking** model — answers may include a reasoning
|
|||
|
|
block; that is expected and improves quality.
|
|||
|
|
|
|||
|
|
### 6. Quick API smoke test
|
|||
|
|
|
|||
|
|
```bash
|
|||
|
|
curl -s http://localhost:8081/v1/chat/completions \
|
|||
|
|
-H "Content-Type: application/json" \
|
|||
|
|
-H "Authorization: Bearer sk-local-bonsai" \
|
|||
|
|
-d '{
|
|||
|
|
"model": "local",
|
|||
|
|
"messages": [{"role":"user","content":"Say hello in one short sentence."}],
|
|||
|
|
"temperature": 0.7,
|
|||
|
|
"max_tokens": 128
|
|||
|
|
}'
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
(`model` can be any string or the id returned by `/v1/models` — llama-server is
|
|||
|
|
usually lenient.)
|
|||
|
|
|
|||
|
|
### 7. What not to do
|
|||
|
|
|
|||
|
|
- Do **not** route bonsai/open-webui through gluetun — inference is local and
|
|||
|
|
the VPN killswitch would only add pain.
|
|||
|
|
- Do **not** expect stock Ollama to load `Ternary-Bonsai-27B-Q2_0.gguf` without
|
|||
|
|
Prism’s kernels (use our `bonsai` service).
|
|||
|
|
- Keep Open WebUI / OpenHands off the public internet; use LAN + auth (`WEBUI_AUTH=true`).
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Layout after AI setup
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
arr-stack/
|
|||
|
|
├── docker-compose.yml
|
|||
|
|
├── .env / .env.example
|
|||
|
|
├── ai/ # Prism llama-server image
|
|||
|
|
│ ├── Dockerfile
|
|||
|
|
│ └── entrypoint.sh
|
|||
|
|
├── models/bonsai/ # GGUF weights (gitignored content)
|
|||
|
|
├── workspace/ # files for RAG / OpenHands
|
|||
|
|
├── searxng/settings.yml
|
|||
|
|
├── scripts/download-bonsai-model.sh
|
|||
|
|
├── config/ # app configs (open-webui, sonarr, …)
|
|||
|
|
├── data/ # media + torrents
|
|||
|
|
└── wireguard/wg0.conf
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
## Notes
|
|||
|
|
|
|||
|
|
- All app-to-app traffic uses container DNS names (`gluetun`, `sonarr`,
|
|||
|
|
`bonsai`, …) on the internal `arr-net` bridge; only WebUIs are exposed to
|
|||
|
|
your LAN.
|
|||
|
|
- Update media stack with `docker compose pull && docker compose up -d`.
|
|||
|
|
Gluetun is pinned to major version `v3` — read its release notes before
|
|||
|
|
jumping majors.
|
|||
|
|
- Rebuild AI runtime after Prism release bumps: edit `PRISM_RELEASE` in
|
|||
|
|
`ai/Dockerfile`, then `docker compose build --no-cache bonsai`.
|
|||
|
|
- Optional media extras (commented in compose): Jellyseerr.
|
|||
|
|
- Optional coding profile: `docker compose --profile coding up -d openhands`.
|