RackNerd VPS runbook for Shouon Al-Ghithaa — SSH from zero to multi-project ops: adding projects, staging/production isolation, zrok tunnels, domains, backups, and plan upgrades. Written for a firs...

RackNerd Server Runbook — Shouon Al-Ghithaa

Audience: someone who has never used SSH or administered an Ubuntu server, and needs to operate this specific server safely. Also written for Claude: every command is tagged, every invariant is explicit. If you are an AI agent, read §0 and §15 before running anything.


§0 — How to read this document

Command tags

Every command block carries one of these. Never run a WRITES command you do not understand.

Tag Meaning
[READ-ONLY] Changes nothing. Safe to run any time, any environment. Run these freely.
[WRITES] Changes files, services, or data. Reversible unless stated.
[DESTRUCTIVE] No undo. Data loss possible. Take a backup first and read the whole section.

The five rules that matter more than any command here

  1. The copy decides the environment. There is no --production / --staging flag anywhere. Every tool reads the deploy/.env file sitting next to it:
    /opt/shouon          → database "shouon"          → PRODUCTION
    /opt/shouon-staging  → database "shouon-staging"  → STAGING
    
    So cd /opt/shouon-staging && ./deploy/shouonctl restore cannot touch production. A flag can be forgotten after you press Enter; a path is visible before you press it.
  2. Always ask the server which environment you are in, before any command that writes. It is one read-only command (§4.1). user add does not ask you — and it once wrote an admin account into the production database from the staging copy.
  3. shouonctl update never runs SQL. By design. It updates code, services and web config. It does not create a table, alter a column, or run a migration. When a database file changed in a pull, update prints a warning line at the end of its output naming those files. That warning line is the source of truth for "what needs running" — not git diff, not this document.
  4. Run update and check as a pair. update deploys; check measures. A green check is the only statement about the server's health that means anything.
  5. "Not measured" is not "broken", and it is not "fine" either. check says ! when it could not measure something (no data to probe with, for example). That is an honest third answer. Do not read it as a failure, and do not read it as a pass.

§1 — This server at a glance

All numbers below were measured on the server, not assumed.

Provider RackNerd (KVM VPS)
IP 204.44.93.211
OS Ubuntu 24.04 LTS
RAM 1.9 GB (~500 MB in use)
Disk ~34 GB (~13% used)
Swap none configured — see §14.6, this matters
Login root over SSH on port 22
Open to the internet 22, 80, 443 only (ufw: deny incoming by default)
Everything else bound to 127.0.0.1 — unreachable from outside

What runs on it

Service Port Role
postgresql 5432 the database
postgrest 3000 auto-generated REST API over the database
shouon-auth 9999 login/session service + live event stream (SSE)
shouon-storage 9000 file uploads
caddy 80, 443, 8088 web server + automatic HTTPS certificates
php8.3-fpm — the /api/*.php AI endpoints
coturn — relay for voice/video calls behind restrictive networks

Public entry points

URL Serves Note
https://shouon-al-ghethaa.com production the paid domain; clients use this
https://www.shouon-al-ghethaa.com → redirects 301 only, never serves (§8.4)
https://204.44.93.211.sslip.io production safety net; saved us during a 12h tunnel outage
https://shouonalghithaa.share.zrok.io staging zrok tunnel, retargeted 2026-09-29
⚠️ The zrok link used to serve production and was given to clients. It now serves staging (an empty database). If a client reports "my account is gone", that is why — send them to the paid domain.

Where things live on disk

Path What
/opt/shouon production code (a git clone)
/opt/shouon-staging staging code (a separate git clone)
/opt/shouon/deploy/.env the secrets file. Not in git. Losing it is bad (§11.4)
/etc/shouon/ service config: postgrest.conf, auth.env, storage.env
/var/lib/shouon/storage/ uploaded files
/var/backups/shouon/ database backups
/etc/caddy/conf.d/shouon.caddy this project's web config
/etc/caddy/Caddyfile server-wide. Do not edit — regenerated on every update

§2 — SSH from absolute zero

SSH is a text connection to the server. You type a command, the server runs it and prints the result. There is no mouse and no undo.

2.1 Connect

Windows 10/11 — open PowerShell (Start → type powershell):

ssh root@204.44.93.211

macOS / Linux — open Terminal:

ssh root@204.44.93.211

First time only, it asks:

The authenticity of host '204.44.93.211' can't be established.
ED25519 key fingerprint is SHA256:...
Are you sure you want to continue connecting (yes/no)?

Type yes and press Enter. This happens once per computer. It is the server introducing itself — not an error.

Then it asks for the password. The screen shows nothing while you type. No dots, no stars. That is normal, not a frozen terminal. Type it and press Enter.

Success looks like:

root@racknerd-xxxxx:~#

2.2 Reading the prompt

root@racknerd-xxxxx:~#
└┬─┘ └──────┬─────┘ │ │
 │          │       │ └── # means root: no command will ask "are you sure?"
 │          │       └──── current directory (~ = /root)
 │          └──────────── the server's name
 └─────────────────────── you are root (full power, no safety net)

2.3 The eight commands you actually need

# [READ-ONLY] where am I?
pwd

# [READ-ONLY] what is in here?
ls -la

# [READ-ONLY] go to the project
cd /opt/shouon

# [READ-ONLY] read a file, one screen at a time (q to quit)
less /etc/caddy/conf.d/shouon.caddy

# [READ-ONLY] is a service alive?
systemctl is-active postgresql

# [READ-ONLY] why did a service fail? (last 50 lines)
journalctl -u shouon-auth -n 50 --no-pager

# [READ-ONLY] disk and memory
df -h / && free -h

# leave the server
exit

2.4 Things that will confuse you once

  • A command printed nothing. On Unix, success is usually silent. No output is good news.
  • The terminal is stuck. Press Ctrl+C to cancel the running command. If you are inside a file viewer, press q.
  • You are inside a text editor and cannot escape. In nano: Ctrl+X, then N to discard. In vim: press Esc, then type :q! and Enter.
  • shouonctl logs appears to hang. It is a live log follower — it waits for new lines forever. Press Ctrl+C. To search logs instead, see §12.2.
  • Arabic text looks like ????. Your terminal font/encoding. The server is fine; use Windows Terminal or iTerm2.

2.5 Stop using a password — use a key (recommended, 3 minutes)

A password can be brute-forced; fail2ban is installed but a key is strictly better.

On your own computer [WRITES - your computer only]:

ssh-keygen -t ed25519 -C "my-laptop"
# press Enter 3 times to accept defaults

Copy it to the server:

# macOS / Linux
ssh-copy-id root@204.44.93.211

# Windows PowerShell (no ssh-copy-id):
type $env:USERPROFILE\.ssh\id_ed25519.pub | ssh root@204.44.93.211 "mkdir -p ~/.ssh && cat >> ~/.ssh/authorized_keys"

Now ssh root@204.44.93.211 logs in with no password.

⚠️ Verify the key works in a second terminal window before disabling password login. Disabling passwords while your key is broken locks you out of your own server, and the only way back in is RackNerd's web console (§14.2).

§3 — The mental model: how one server holds many projects

3.1 The principle: no shared file is ever edited

Resource Shared Per project
Caddy /etc/caddy/Caddyfile — never edited /etc/caddy/conf.d/.caddy
PostgreSQL the service one database + roles per project
Local port — one per project
systemd units — names prefixed with the project name
zrok one account, one environment (= this machine) one share per project

Adding a project = creating its files. Removing it = deleting them. You never open a file that serves something else.

3.2 The request path

Browser
   │  https://shouon-al-ghethaa.com
   ▼
Caddy :443 ──reads the "Host" header──┐
   │                                   │
   │                      ┌────────────┴─────────────┬──────────────────┐
   ▼                      ▼                          ▼                  ▼
conf.d/shouon.caddy   conf.d/shouon-staging    conf.d/project2    (unknown name)
   │                      .caddy                  .caddy               │
   │                                                                   ▼
   ├── /            → static files from /opt/shouon (index.html)   no certificate
   ├── /rest/v1/*   → 127.0.0.1:3000  (PostgREST)                  → TLS fails
   ├── /auth/v1/*   → 127.0.0.1:9999  (auth + SSE)                 ERR_SSL_
   ├── /storage/v1/*→ 127.0.0.1:9000  (files)                      PROTOCOL_ERROR
   └── /api/*.php   → php8.3-fpm

3.3 Port registry — reserve before you use

Port Owner
3000 / 9999 / 9000 Shouon production: REST / auth / storage
8088 Shouon production — local tunnel entry
3001 / 9998 / 9001 Shouon staging
8089 Shouon staging — local tunnel entry
8090 and up next projects
5432 PostgreSQL (all projects share the service)
2019 Caddy admin API

[READ-ONLY] What is actually listening right now:

ss -tlnp | awk '{print $4}' | grep -oE '[0-9]+
  
  




 | sort -un
🔴 Two projects on one port do not both fail — one fails silently and the other works. Caddy rejects a duplicate site name with ambiguous site definition, but a duplicate backend port is worse: one service refuses to start with nothing obvious in the output, the other keeps serving, and later someone reports that a project "broke for no reason."

3.4 What is NOT separated between environments — stated plainly

postgres, authenticator and auth_service are cluster-level PostgreSQL roles, not per-database. Generating new passwords for a second environment would run ALTER ROLE and overwrite production's password — production would fail instantly with password authentication failed.

So those three passwords are copied as-is, and what gets generated fresh is JWT_SECRET, ANON_KEY and PASSWORD_VAULT_KEY. That is the separation that actually matters: a session token minted in staging is rejected by production.

⚠️ The price, said honestly: whoever reads the staging config file can connect to the production database. This adds no exposure — both files are mode 600, same owner, same machine; anyone who can read one can read the other.


§4 — Daily operations: shouonctl

shouonctl is the single tool. Run it from inside the environment you mean:

cd /opt/shouon && ./deploy/shouonctl       # production
cd /opt/shouon-staging && ./deploy/shouonctl    # staging
The bare shouonctl (no path) is a symlink to the production copy. If you are not sure which copy you are about to drive, use ./deploy/shouonctl with an explicit cd.

4.1 Ask which environment you are in — first, every time

[READ-ONLY]

cd /opt/shouon && ./deploy/shouonctl env

It prints the project name, database name, copy path, ports, and the public URL. This touches no service and no data. Run it before anything that writes.

4.2 The command table

Command Tag What it does
env [READ-ONLY] which environment / database / copy am I?
status [READ-ONLY] all services: state, memory, uptime
monit [READ-ONLY] same, live, refreshes every 2s (Ctrl+C to exit)
check [READ-ONLY] the full audit: security, certificate, backups, resources, RLS probes
logs [service] [READ-ONLY] live log follower (Ctrl+C to exit)
errors [hours] [READ-ONLY] errors only, default last hour
tunnel [READ-ONLY] which environment does the zrok tunnel point at?
tunnel here [WRITES] retarget the tunnel to this environment
restart [service] [WRITES] no name = all, in the correct order
stop / start [WRITES] —
update [WRITES] backup → git pull → re-stamp frontend → sync config → restart
backup [WRITES] take a backup now
user list [READ-ONLY] list accounts
user add|passwd|role|del [WRITES] manage accounts — does not ask which environment
psql [WRITES] raw database shell. No undo. Read §11.3 first
restore [DESTRUCTIVE] restore a backup over the current database
restore-from-supabase [DESTRUCTIVE] import an old Supabase project
bootstrap [DESTRUCTIVE] build the schema from scratch
sync-refs --from [READ-ONLY] without --apply copy reference rows into this environment
repair-storage [READ-ONLY] without --apply strip multipart wrappers from uploaded files
migrate-storage [READ-ONLY] without --apply pull remaining Supabase files onto this server
enable-db-download [WRITES] once admin "download full backup" button
enable-password-vault [WRITES] once admin "show current password" feature
enable-voice-calls [WRITES] once installs coturn, opens its ports

4.3 The normal deploy

[WRITES]

cd /opt/shouon
sudo ./deploy/shouonctl update
sudo ./deploy/shouonctl check

Always send/run both. update deploys, check measures.

What update touches and what it does not:

Updated frontend code · services · Caddy config · systemd units
Never touched the database — no table, no row · uploaded files · .env · user accounts
🔴 Never run a bare git pull in /opt/shouon. update strips the deploy-time stamps from index.html (server URL, anon key, version) before pulling and rewrites them after. A plain git pull hits Your local changes would be overwritten — and if you force past it, the app ends up pointing at the old Supabase URL.

4.4 When update says database files changed

At the very end of its output you may see:

⚠️ Database files changed in this pull — and update does not run SQL:
   security/16-quality.sql
   deploy/postgres/init/04-schema-compat.sql

That list is the instruction. Nothing happened to the database yet.

[WRITES] — run them yourself, in the order printed, one at a time:

cd /opt/shouon
./deploy/shouonctl env          # confirm the database name first
sudo -u postgres psql -d shouon -f security/16-quality.sql

Each of these files prints its own self-verification (✓ measured: ...) at the end and refuses to claim success it did not measure. Read that line. Then:

sudo ./deploy/shouonctl check
🔴 Deploy before SQL, always. The file only exists on the server after update pulls it. Running psql -f security/16-quality.sql first gives No such file or directory — this happened in production. 🔴 And the warning line is the source of truth, not the repository. A hand-written list of files taken from GitHub once missed five files that had changed on the server.

§5 — Scenario A: add a brand-new project (a different app)

Use this when the new thing is not Shouon — a different codebase, different data. (For a second environment of Shouon itself, use §6 instead; it is automated.)

Step 0 — Decide and reserve, on paper, before touching anything

Decide Example
project name (lowercase, letters/digits/hyphen) project2
local tunnel port 8090 (next free — §3.3)
backend ports if it has an API 3002, 9997, 9002
database name project2
public address a domain (§8) or a zrok link (§7)

[READ-ONLY] Verify every port is actually free:

ss -tlnp | awk '{print $4}' | grep -oE '[0-9]+
  
  




 | sort -un

[READ-ONLY] Verify you have the RAM to spare:

free -h && systemd-cgtop -n1 --order=memory | head -12

Shouon production is ~470 MB of 1.9 GB. If the remainder is thin, read §14 before adding a project — the kernel kills a random process when memory runs out, and it is usually PostgreSQL.

Step 1 — Put the code on the server

[WRITES]

cd /opt
git clone  project2

Step 2 — Its own database and role

[WRITES]

sudo -u postgres createdb --encoding=UTF8 --lc-collate=C --lc-ctype=C \
     --template=template0 project2
sudo -u postgres psql -c "CREATE ROLE project2_app LOGIN PASSWORD 'a-strong-password'"
sudo -u postgres psql -c "GRANT ALL ON DATABASE project2 TO project2_app"
One database per project, not one schema inside a shared database: a backup of one project must not carry another project's data, and a wrong restore must not take down everything.

Step 3 — Its own Caddy file

[WRITES]

nano /etc/caddy/conf.d/project2.caddy
# Public site — only if you own a domain for it (see §8)
project2.example.com {
    encode zstd gzip
    header Strict-Transport-Security "max-age=31536000; includeSubDomains"
    root * /opt/project2
    file_server
}

# Tunnel entry: a port with NO site name, so it matches any Host header.
# bind 127.0.0.1 keeps it unreachable from the internet.
:8090 {
    bind 127.0.0.1
    encode zstd gzip
    root * /opt/project2
    file_server
}
🔴 Write :8090, not http://127.0.0.1:8090. Caddy reads the second form as a domain name, so no request ever matches it and the tunnel returns a bare 404 — "the link opens and gives 404" with nothing in the logs. 🔴 If this project defines a reusable Caddy snippet, its name must be unique across all of conf.d. Snippet names are global, not local to their file. Two files defining (shouon_app) make Caddy reject the entire configuration — so adding a project takes production down with it, not just itself.

Step 4 — Validate BEFORE reloading

[READ-ONLY]

caddy validate --config /etc/caddy/Caddyfile --adapter caddyfile
🔴 validate before reload, never after. Caddy loads conf.d/* as one configuration. A broken new file that gets applied takes down every project on the server, not only the new one.

Only if validation passed, [WRITES]:

systemctl reload caddy

Step 5 — Verify end to end

[READ-ONLY]

# the local tunnel entry must serve the page
curl -s -o /dev/null -w '%{http_code} %{size_download}\n' \
     -H 'Host: probe.example' http://127.0.0.1:8090/

# existing projects must be untouched
cd /opt/shouon && ./deploy/shouonctl check

Expect 200 and a byte count that looks like a real page. 200 with a tiny size is a failure, not a success — something else answered, not your project.


§6 — Scenario B: add a staging environment for Shouon

This is automated by one script with a preflight that refuses on any collision before writing anything.

6.1 What you get

Production Staging
Copy /opt/shouon /opt/shouon-staging
Database shouon shouon-staging
Config /etc/shouon /etc/shouon-staging
Storage /var/lib/shouon/storage /var/lib/shouon-staging/storage
Backups /var/backups/shouon /var/backups/shouon-staging
Units postgrest, shouon-auth, shouon-storage shouon-staging-*
Ports 3000 / 9999 / 9000 / 8088 3001 / 9998 / 9001 / 8089
Caddy snippet (shouon_app) (shouon_staging_app)
JWT_SECRET its own its own and different

6.2 Create it

[WRITES] — run it from the production copy, as root:

cd /opt/shouon
sudo ./deploy/scripts/new-env.sh \
     --name shouon-staging \
     --dir /opt/shouon-staging \
     --public-url https://shouonalghithaa.share.zrok.io \
     --local-port 8089 --rest-port 3001 --auth-port 9998 --storage-port 9001

What it does, in order:

  1. Preflight — refuses if the directory, database, any port, any systemd unit, or the Caddy site already exists. Nothing is written if it refuses.
  2. Takes a backup of the current environment first.
  3. Clones the code from the local copy (same commit, works offline), then points origin back at the real remote so update works later.
  4. Writes a new .env: copies the three cluster role passwords, generates a new JWT_SECRET, ANON_KEY and PASSWORD_VAULT_KEY, and verifies they differ from production's — it aborts if they do not.
  5. Creates the database and copies the schema from the live database with zero rows.
  6. Writes /etc/shouon-staging/*, installs Node dependencies, renders and installs systemd units.
  7. Applies the role/compat SQL layer (01, 03, 04).
  8. Renders the Caddy site, validates the whole conf.d directory together, and only then installs and reloads.
  9. Starts the services and measures: the tunnel entry serves the page, PostgREST answers, and — the decisive check — reads /proc//environ of the live auth process to confirm it is connected to the new database. It dies if it is not.
🔑 Why the schema is copied from the live database, not built from the repository: building from the repo gives ~57 tables; the production database has 193 — it came from a Supabase restore, not from our schema files. A test environment built from the repo is therefore not a copy of production: a bug can appear in one and not the other. Pass --bootstrap only if you explicitly want the repo's schema instead.

6.3 What is left for you afterwards

[WRITES] Create an admin account in staging (note the cd):

cd /opt/shouon-staging
sudo ./deploy/shouonctl user add

[READ-ONLY] Confirm you really are in staging and that production is intact:

cd /opt/shouon-staging && ./deploy/shouonctl env && ./deploy/shouonctl user list
cd /opt/shouon         && ./deploy/shouonctl env && ./deploy/shouonctl user list

Production should list its real accounts; staging should list only the one you just created. That is the measurement that proves the two are separated — more convincing than reading a config line.

[WRITES] Install the staging environment's own scheduled jobs (backups):

cd /opt/shouon-staging
sudo ./deploy/scripts/cron-install.sh
Run production's first if neither has been installed since the fix, then staging's. The cron file is named after the project (/etc/cron.d/shouon-staging), plus one shared machine-level file (/etc/cron.d/shouon-host) for disk monitoring.

6.4 Staging has no public domain — on purpose

Staging is served only through the zrok tunnel, so its .env has DOMAIN=, DOMAIN_SERVE= and DOMAIN_REDIRECT= deliberately empty, and Caddy renders only the :8089 tunnel entry.

🔴 If you wrote a site block for the zrok hostname, Caddy would request a Let's Encrypt certificate for a name we do not control — it cannot work (the tunnel provider terminates TLS with its own certificate), and failed attempts are rate-limited per name for hours, which blocks your next attempt for a reason unrelated to your configuration.

That is why PUBLIC_URL is a separate key from DOMAIN: the first is stamped into the frontend, the second is routed to. Staging has the first and not the second.

6.5 Empty reference tables in staging are the design, not a bug

The schema was copied with zero rows, so dropdowns (departments, branches, quality items) are empty and the app can look broken.

[READ-ONLY] first — it reads and writes nothing without --apply:

cd /opt/shouon-staging
sudo ./deploy/shouonctl sync-refs --from shouon departments

[WRITES] once the counts printed look right:

sudo ./deploy/shouonctl sync-refs --from shouon departments --apply

The destination is this copy's database and only the source is passed, so the transfer cannot run backwards. It refuses to write over existing rows, counts missing foreign-key targets up front, and verifies the ID sequence advanced afterwards.

⚠️ branches carries manager_id → users, and staging has one user. Run the read-only pass first: it tells you exactly how many references are missing instead of failing halfway through.

§7 — Scenario C: zrok tunnels, end to end

zrok gives a public https://.share.zrok.io URL to a port bound to 127.0.0.1, with no open inbound port and no certificate of our own. The connection is outbound from our server to zrok's edge.

7.1 The naming model — get this right first

your zrok account
└── racknerd-vps              ← ZROK_ENVIRONMENT_NAME = THE MACHINE (set once)
    ├── shouonalghithaa       ← ZROK_UNIQUE_NAME → shouonalghithaa.share.zrok.io
    └── project2              ← ZROK_UNIQUE_NAME → project2.share.zrok.io
  • 🔴 Never name the environment after a project (e.g. shouon-vps). The environment is the machine, and it will hold other projects; the name becomes a lie to whoever reads the tree in six months.
  • ⚠️ ZROK_UNIQUE_NAME is global across all zrok users — it is a shared subdomain. A taken name is rejected at reservation time. 4–32 chars, lowercase letters and digits.

7.2 Which environment does the tunnel currently serve?

[READ-ONLY]

cd /opt/shouon && ./deploy/shouonctl tunnel

It reads the port out of the systemd unit's own env file (not a guessed path) and tells you whether it points at this environment or another one.

7.3 Retarget the tunnel to the current environment

[WRITES]

cd /opt/shouon-staging
sudo ./deploy/shouonctl tunnel here

It is safe by construction:

  1. Gate before writing — it refuses if this environment's local entry does not already serve the page. Pointing the tunnel at a dead port disables the clients' link completely until someone notices, which is worse than not switching.
  2. Backs up the env file with a timestamp.
  3. Rewrites ZROK_TARGET, then reads it back — sed with no match succeeds silently, so "wrote it" must be measured, not assumed.
  4. Restarts the service; if it does not come up, restores the backup and restarts again.

The manual equivalent, if you ever need it [WRITES]:

sed -i 's|^ZROK_TARGET=.*|ZROK_TARGET="http://127.0.0.1:8089"|' \
  /opt/openziti/etc/zrok/zrok-share.env
grep -nE '^[A-Z_]+=.*#' /opt/openziti/etc/zrok/zrok-share.env   # must print nothing
systemctl restart zrok-share
🔴 This is a decision to state before executing, not after. The zrok link was given to clients. Pointing it at staging means a client who opens it sees an empty database and cannot find their account. It is only acceptable because the paid domain is now measured to be serving, with sslip.io as a backup.

7.4 Add a second zrok share for a second project

The zrok-share package ships one service. For the second, copy the unit and its env file [WRITES]:

cp /opt/openziti/etc/zrok/zrok-share.env /opt/openziti/etc/zrok/project2.env
nano /opt/openziti/etc/zrok/project2.env
#   → set ZROK_UNIQUE_NAME to the new name
#   → set ZROK_TARGET="http://127.0.0.1:8090"
#   → DO NOT change ZROK_ENVIRONMENT_NAME — one environment per machine

sed 's|zrok-share.env|project2.env|g' /lib/systemd/system/zrok-share.service \
  > /etc/systemd/system/zrok-project2.service
systemctl daemon-reload && systemctl enable --now zrok-project2

7.5 🔴 Never leave a trailing comment on a value line

The shipped zrok-share.env contains example comments at line ends:

ZROK_TARGET="http://127.0.0.1:8088"                  # e.g., http://127.0.0.1:3000

The service script parses the file itself — it does not source it in a shell. So it does not strip the comment and does not understand quotes. The effective target becomes:

'http://127.0.0.1:8088# e.g., http://127.0.0.1:3000'

which contains spaces, so reservation fails with Error: accepts between 1 and 2 arg(s), received 4. This happened on the first run.

[READ-ONLY] after editing — must print nothing:

grep -nE '^[A-Z_]+=.*#' /opt/openziti/etc/zrok/zrok-share.env

7.6 🔴 The real state directory, and a service that restarts every 3 seconds

The zrok-share unit runs with DynamicUser, so /var/lib/zrok-share is only a link to /var/lib/private/zrok-share. The unit restarts every 3 seconds on failure — so deleting the state directory while the service is looping recreates it immediately and you chase your own tail.

[WRITES] Stop it first, always:

systemctl stop zrok-share
ls -la /var/lib/private/zrok-share/.zrok/

7.7 🔴 409 shareConflict — a reservation stranded on zrok's side

If reservation succeeded on zrok's servers but reserved.json was then lost locally (a failure mid-way, or wiping state while the service looped), the name stays reserved there and unknown here. Every attempt after that:

[ERROR]: unable to create share (... [POST /share][409] shareConflict)
The rule: zrok release BEFORE deleting the state directory, never after. Deleting state burns the share — it stays reserved on zrok's server and nobody owns it. zrok release cannot fix it afterwards, because release runs as the caller's environment identity, and that identity lived inside the identities/ directory you just deleted. The service replies [DELETE /unshare][404] unshareNotFound, and no CLI command can delete another environment's share.

The recovery that actually worked on this server:

# 1) [WRITES] disable the current environment (re-registers later from the token)
systemctl stop zrok-share
HOME=/var/lib/private/zrok-share zrok disable

# 2) In the zrok web console: open the share → red delete button.
#    FIRST confirm Share Token, Env Z Id and Target — do not delete someone
#    else's share. Then delete the leftover empty racknerd-vps environment.

# 3) [WRITES] register a fresh environment and reserve the name again
systemctl start zrok-share
Do not pick a new name to dodge a 409. Every name you burn stays reserved on your account, and you will run out of names that mean anything.

7.8 🔴 The tunnel blocks in.(…) queries — and tells no one

Measured in production (2026-09-21):

Request Via tunnel Direct domain Caddy locally
?user_id=eq.999999 ✅ 401 ✅ 401 ✅ 401
…&section_key=in.(a,b,c) ❌ 403 ✅ 401 ✅ 401
…&section_key=in.%28a,b,c%29 ❌ 403 ✅ 401 ✅ 401

A security filter at the zrok provider reads in.(a,b,c) in the query string, mistakes it for SQL injection, and returns its own 403 HTML page before the request reaches us. Percent-encoding does not help — %28 is decoded, then matched.

Why it went unnoticed: ordinary reads use eq. and pass. in.(…) appears in 52 queries, mostly dashboards that come back empty when blocked — so they look like "no data" rather than "broken."

The fix lives on both ends we control, and must stay in sync:

Where What
index.html → escapeApiParens in.(a,b,c) → in.!lp!a,b,c!rp! before sending
Caddy shared snippet uri replace "!lp!" "(" … four rules
🔴 The only real danger is the two ends drifting apart. A marker changed on one side and forgotten on the other does not break loudly — !lp! arrives at PostgREST literally and it replies with a vague error about a column that does not exist. A guard (deploy/test/paren-escape.mjs) extracts the markers from the frontend and demands a Caddy rule for each, and check.sh measures a disguised path through the live tunnel.

This still applies to staging after retargeting — the filter is at the provider, whatever sits behind the tunnel. The Caddy template is shared, so the rules render for both environments.

7.9 When the tunnel is down and it is not our fault

Measured, twice:

  • 2026-09-26 — Public Shares: Major Outage on zrok's status page while API / Console was Operational. Our server reached api.zrok.io fine (200 in 0.2 s) but the share did not answer. It recovered on its own after ~12 hours, with nothing touched on our side.
  • 2026-09-27 — the tunnel returned 403 with header server: awselb/2.0 — i.e. from zrok's AWS edge, not from Caddy — while the link worked from the client's browser. The edge was refusing our server's own outbound probe only.

[READ-ONLY] How to tell "zrok is broken" from "we are broken":

# does OUR entry serve the page, bypassing the tunnel entirely?
curl -s -o /dev/null -w '%{http_code} %{size_download}\n' \
     -H 'Host: shouonalghithaa.share.zrok.io' http://127.0.0.1:8088/

# the tunnel, as a browser would, following compression
curl -s --compressed -o /dev/null \
     -w '%{http_code} %{time_total}s %{size_download}\n' --max-time 60 \
     https://shouonalghithaa.share.zrok.io/

# did the request reach US? our snippet adds this header; zrok's edge does not
curl -sSI https://shouonalghithaa.share.zrok.io/ | grep -i 'x-content-type-options\|^server'
🔑 The status code alone does not tell you who answered. A 403 from a security filter looks exactly like a 403 from an application. Only a header distinguishes them: Caddy's shared snippet adds nosniff and removes the Server header; zrok's edge does the opposite. This cost a full round of diagnosis once, when a 403 from a proxy in a different environment was read as an answer from the destination.

🔴 Do not wipe the state directory and do not re-register during an outage. The state was fine both times; wiping it burns the share (§7.7) and fixes nothing. The remedy is to wait, and to send clients to the direct domain.

§8 — Scenario D: attach a domain to a project

8.1 The misconception worth killing first

A DNS record does not point at a project. It points at the server. Every domain on this server — however many projects it hosts — has the same A record and the same IP.

shouon-al-ghethaa.com   ─┐
project2.example.com    ─┼──►  A  204.44.93.211   (one identical record)
hr.othercompany.sa      ─┘
                                      │
                    Caddy reads the Host header of the request:
                                      │
      ┌───────────────────────────────┼──────────────────────────┐
      ▼                               ▼                          ▼
 conf.d/shouon.caddy            conf.d/project2.caddy      conf.d/hr.caddy

Separation happens in the site name written inside the Caddy block, not in the network. A request for a name no file in conf.d knows reaches no project: Caddy has no certificate for it, TLS fails, and the browser shows ERR_SSL_PROTOCOL_ERROR.

🔑 That exact error is proof the DNS worked. A missing or wrong record gives DNS_PROBE_FINISHED_NXDOMAIN or a timeout. A TLS failure means the name resolved, TCP opened on our port 443, and only the handshake failed — so the rest is entirely on our side.

8.2 What to ask the domain owner for

Type Name Value TTL
A @ the server IP 300
A www the server IP 300

And nothing else. No CNAME, no TXT, no nameserver change, no port to open. The certificate issues itself from Let's Encrypt over port 80.

  • ⚠️ TTL 300 is deliberate for the first two days: a mistake is fixed in five minutes instead of twenty-four hours. Raise it after it settles.
  • 🔴 Ask them to delete any existing A or CNAME on @ or www — the registrar's "this domain is for sale" parking page. Leaving it does not block the site cleanly; it makes the site work once and fail once depending on which record the browser got, which is worse than not working at all because it looks like a random intermittent fault.

8.3 The order — and why it is not a detail

1) A records are added           ← at the domain owner
2) propagation is MEASURED        ← from the server
3) deploy/.env is set             ← then shouonctl update

[READ-ONLY] Step 2, from the server:

getent ahostsv4 shouon-al-ghethaa.com
getent ahostsv4 www.shouon-al-ghethaa.com

Both must print the server's IP. If they do not, stop and wait.

🔴 Measure before step 3, not after. If Caddy requests a certificate while the record has not propagated, it fails — and Let's Encrypt rate-limits failed attempts for hours on that name. Your next attempt is then blocked for a reason that has nothing to do with your configuration.

8.4 Configure it — three keys, nothing more

[WRITES]

cd /opt/shouon
nano deploy/.env
DOMAIN=shouon-al-ghethaa.com                  # primary: serves, gets HSTS
DOMAIN_SERVE=204.44.93.211.sslip.io           # also serves (space-separated list)
DOMAIN_REDIRECT=www.shouon-al-ghethaa.com     # 301s onto the primary, never serves
ACME_EMAIL=you@example.com                    # certificate-expiry notices

[WRITES] Render, validate, reload, restart, measure — all inside update:

sudo ./deploy/shouonctl update
sudo ./deploy/shouonctl check
🔴 www must redirect, never serve — this is not cosmetic. The frontend builds its API base from location.origin at runtime, and a browser separates session, storage and service worker per origin. If www served the app itself, a user who visits www.x.com and then x.com is "not logged in" for no visible reason, with two caches fighting. Redirecting makes the origin singular. ⚠️ Keep the old link in DOMAIN_SERVE, do not drop it. On 2026-09-26 the zrok tunnel was down for twelve hours and sslip.io was the clients' only way in. Dropping it removes a safety net that has actually been used. (Its cost is a second origin with a separate session — acceptable because nobody reaches it by accident.) ✅ The zrok tunnel is unaffected by any of this. Its entry is a :8088 block with no site name, so it matches any Host. Changing the domain or adding ten does not touch it.

8.5 Verify from all four entry points

[READ-ONLY]

for u in https://shouon-al-ghethaa.com \
         https://www.shouon-al-ghethaa.com \
         https://204.44.93.211.sslip.io \
         https://shouonalghithaa.share.zrok.io ; do
  printf '%-45s ' "$u"
  curl -s -o /dev/null -w '%{http_code} %{size_download}B %{time_total}s\n' \
       --compressed --max-time 30 "$u"
done

# certificate: subject and expiry
echo | openssl s_client -connect shouon-al-ghethaa.com:443 \
       -servername shouon-al-ghethaa.com 2>/dev/null \
     | openssl x509 -noout -subject -enddate

What a correct result looks like (measured 2026-09-28):

URL Expect
root domain 200, ~827,398 B
www 301 — not 200
sslip.io 200, the same byte count
zrok 200 (whichever environment it targets)
🔑 Identical byte counts across entry points is a cross-check that happens for free: it proves they serve the same file from the same server, not a stale copy behind one of them. If one differs, that entry point is serving something else.

8.6 Does a domain block adding more projects later? — No

Measured: the server-wide Caddyfile only does import conf.d/*.caddy with no name of its own, and every project owns its deploy/.env, its conf.d file, its port and its database. update validates the whole conf.d directory together, so a conflict with another project is caught before it is applied. The only machine-wide value is ACME_EMAIL (who gets certificate-expiry mail), which does not affect routing.


§9 — Scenario E: move clients from the zrok link to a paid domain

The sequence matters, because the zrok link is already in clients' hands.

  1. [READ-ONLY] Confirm the paid domain serves correctly — §8.5. Do not proceed on a curl that was not run.
  2. Tell the clients the new address, and keep both working for a grace period. Both point at the same production database, so nothing is lost either way.
  3. [READ-ONLY] Only once the domain is confirmed, decide whether the zrok link is still needed for production or can be handed to staging (§7.3).
  4. [WRITES] If handing it to staging: cd /opt/shouon-staging && sudo ./deploy/shouonctl tunnel here.
  5. Announce it: anyone still on the zrok link now sees an empty database.
⚠️ Keep sslip.io in DOMAIN_SERVE throughout. After step 4 the zrok link is no longer a production fallback, so sslip.io is the only one left.

§10 — Scenario F: remove a project or an environment cleanly

[DESTRUCTIVE] — read it all, then take a backup.

# 0) [READ-ONLY] confirm which environment you are about to destroy
cd /opt/shouon-staging && ./deploy/shouonctl env

# 1) [WRITES] final backup, and copy it somewhere off this server
sudo ./deploy/shouonctl backup
ls -lt /var/backups/shouon-staging | head

# 2) [WRITES] stop and disable its services and timers
sudo systemctl disable --now shouon-staging-postgrest shouon-staging-auth \
     shouon-staging-storage
sudo systemctl disable --now shouon-staging-closing-alerts.timer \
     shouon-staging-meeting-alerts.timer shouon-staging-purge-recordings.timer
sudo rm -f /etc/systemd/system/shouon-staging-*.service \
           /etc/systemd/system/shouon-staging-*.timer
sudo systemctl daemon-reload

# 3) [WRITES] remove its Caddy site — validate the rest BEFORE reloading
sudo rm -f /etc/caddy/conf.d/shouon-staging.caddy
caddy validate --config /etc/caddy/Caddyfile --adapter caddyfile   # [READ-ONLY]
sudo systemctl reload caddy

# 4) [WRITES] its scheduled jobs (leave /etc/cron.d/shouon-host alone — shared)
sudo rm -f /etc/cron.d/shouon-staging
sudo systemctl restart cron

# 5) [DESTRUCTIVE] the database — no undo
sudo -u postgres dropdb shouon-staging

# 6) [DESTRUCTIVE] files
sudo rm -rf /opt/shouon-staging /etc/shouon-staging /var/lib/shouon-staging
# keep /var/backups/shouon-staging until you are certain

# 7) [READ-ONLY] prove production survived
cd /opt/shouon && ./deploy/shouonctl check
🔴 Do not delete /etc/cron.d/shouon-host — it is the shared machine-level disk monitor used by every environment. 🔴 Do not delete /var/lib/caddy — it holds the issued certificates. Deleting it forces reissue and can hit the Let's Encrypt weekly limit, taking HTTPS down for all projects.

If the removed environment owned a zrok share, release it before touching any state directory (§7.7).


§11 — Backups, restore, and the gap you should know about

11.1 What runs automatically

Installed per environment by deploy/scripts/cron-install.sh (/etc/cron.d/):

When What
03:00 daily database backup → /var/backups//
04:00 on the 1st restore test of the newest backup — an untested backup is not a backup
hourly (shared file) disk usage alert above 85%

11.2 On demand

[WRITES]

cd /opt/shouon && sudo ./deploy/shouonctl backup

[READ-ONLY]

ls -lt /var/backups/shouon | head

11.3 Restore — the one command with no undo

[DESTRUCTIVE]

cd /opt/shouon
./deploy/shouonctl env          # [READ-ONLY] read this line out loud first
sudo ./deploy/shouonctl backup  # [WRITES] a backup of the CURRENT state
sudo ./deploy/shouonctl restore  /var/backups/shouon/db-YYYY-MM-DD.gz

Restore replaces the live database. The path in your command line is the only thing standing between you and the wrong environment — this is exactly why the environment is chosen by cd and not by a flag.

11.4 🔴 The two gaps, stated plainly

1. Every backup is on the same server. OFFSITE_DEST is not set — a standing decision, but it means losing the server loses the backups too.

[WRITES] To close it, set one line in deploy/.env:

OFFSITE_DEST=s3://your-bucket/shouon          # uses aws cli
# or
OFFSITE_DEST=user@host:/path/to/backups       # uses rsync over ssh

Then the nightly job copies off-server automatically. Verify after the next run: [READ-ONLY]

tail -20 /var/log/shouon-backup.log

2. deploy/.env is not in git and is not in the database backup. It holds JWT_SECRET (whoever has it can mint an admin session token without a password), ANON_KEY, and PASSWORD_VAULT_KEY. Losing it does not lose logins (password hashes are in the database) but does lose the admin "show current password" vault permanently.

[WRITES] Keep a copy somewhere that is not this server:

cp /opt/shouon/deploy/.env /root/shouon-env-backup-$(date +%F)
# then download it to your own machine, from YOUR computer:
#   scp root@204.44.93.211:/root/shouon-env-backup-* .

A copy also exists in /etc/shouon/*.env — which is how a .env loss is recoverable at all, but it is a lost day.

🔴 Never commit .env to the repository.

11.5 Backup-file hygiene

Backup files contain password_plain for every employee. Delete local copies after use. Do not leave them in Downloads, chat, or a shared drive.


§12 — Troubleshooting matrix

12.1 Symptom → likely cause → check

Symptom Likely cause Command
ERR_SSL_PROTOCOL_ERROR in browser DNS resolved but Caddy has no site for that name getent ahostsv4 then §8.4
DNS_PROBE_FINISHED_NXDOMAIN the A record is missing or wrong ask the domain owner
Site works once, fails once an old parking A/CNAME still exists alongside ours §8.2
Certificate will not issue wrong A record, or port 80 closed getent ahostsv4 · ufw status
Could not resolve host during install the server's DNS is broken — unrelated to the repo getent hosts github.com
Permission denied (publickey) on git deploy key not added to the repository §2.5 / repo settings
A service will not start — journalctl -u -n 50 --no-pager
permission denied for table in app the session is not authenticated expected — log in
Login always fails JWT_SECRET differs between services compare grep jwt-secret /etc/shouon/postgrest.conf with /etc/shouon/auth.env
Attachments do not display session cookie not reaching the API frontend and API must be on the same origin
ambiguous site definition from Caddy two conf.d files claim the same name/port §3.3
Caddy rejects the whole config two files define the same snippet name §5 Step 3
Dashboards empty on the zrok link only the in.(…) filter at the provider §7.8
Database stopped accepting writes the disk is full df -h · clean /var/backups
Sudden slowness memory pressure / work_mem × connections §14
A project "broke for no reason" another project took its port and started first §3.3
check prints ! not ✗ it could not measure, not a failure read the line's own wording

12.2 Reading logs correctly

[READ-ONLY]

# live follow — hangs on purpose; Ctrl+C to exit
cd /opt/shouon && ./deploy/shouonctl logs auth

# errors only, last 24 hours, all services
./deploy/shouonctl errors 24
🔴 Never pipe shouonctl logs into grep … | tail. It is a live follower (journalctl -f): tail waits for an end that never comes, the terminal appears frozen, and it looks like "nothing in the log" when it has not read yet. To search instead, read history directly:
journalctl -u shouon-auth --since "-15 min" --no-pager | grep -a "pattern"
-a matters: Arabic lines are sometimes classified as binary and grep goes silent on them.

12.3 Two red lines in check that are decisions, not faults

Line Why
✗ 4 test accounts with known passwords kept by the owner's decision; one is an admin
✗ OFFSITE_DEST not set §11.4 — a standing decision

These are expected to be red. Everything else red is work.


§13 — Hard rules: never do these

Server

  • ❌ Never edit /etc/caddy/Caddyfile — server-wide, regenerated on every update; your edit vanishes silently.
  • ❌ Never reload Caddy before validate — a bad config takes down every project.
  • ❌ Never reuse a port — one service dies silently, the other works (§3.3).
  • ❌ Never share a database between two projects.
  • ❌ Never delete /var/lib/caddy — the certificates live there.
  • ❌ Never delete /etc/cron.d/shouon-host — shared disk monitor.
  • ❌ Never run git clean -xfd in /opt/shouon — it deletes ignored files, which includes .env with JWT_SECRET. Every session drops and the frontend is rebuilt with a key the services do not accept. (git reset --hard and checkout -f do not touch ignored files, so .env survives those.)
  • ❌ Never run a bare git pull in /opt/shouon — use shouonctl update (§4.3).

PostgreSQL

  • ❌ Never set listen_addresses = '*'. To reach the database from your laptop, tunnel it: ssh -L 5432:localhost:5432 root@204.44.93.211
  • ❌ Never raise work_mem above 4 MB. It is per sort operation, not per connection — a query with three sorts takes three times that. Raising it to 16 MB on a 2 GB server is the best-known way to crash it.
  • ❌ Never disable synchronous_commit. This is payroll and accounting data.

Secrets

  • ❌ Never commit .env.
  • ❌ Never paste into chat or a ticket: the Supabase secret key (sb_secret_…), the database connection string, the zrok token, or JWT_SECRET. If a secret does leak, say so explicitly and rotate it.

Process

  • ❌ Never dictate a command you have not read the source of, and never one you have not run yourself. Eight separate production rounds were lost to commands written from assumption (git pull, logs | grep, the .env path, an undefined $DB, a misremembered function name, running SQL before deploying, ! expanding inside double quotes in an interactive shell).
  • ❌ Never claim success you did not measure. "It should work" is not a result.

§14 — Upgrading the RackNerd server to a bigger plan

⚠️ What is measured and what is not, stated up front. Everything in §14.1, §14.5, §14.6 and §14.7 is measured on this server or read from its configuration. §14.2 and §14.3 describe RackNerd's process and must be confirmed with RackNerd support before you pay, because whether an upgrade is in-place or a migration decides whether the IP changes — and that is the whole risk.

14.1 First: measure. Do not buy blind.

[READ-ONLY] Run all of these and write the numbers down:

# memory: the "available" column is what matters, not "free"
free -h

# memory per service, biggest first
systemd-cgtop -n1 --order=memory | head -15

# has the kernel ever killed something for memory? (this is the real signal)
sudo dmesg -T | grep -i -e "out of memory" -e "oom-kill" | tail -20
sudo journalctl --since "-30 days" | grep -ci "out of memory"

# disk: overall, and what is actually large
df -h /
sudo du -xh --max-depth=1 / 2>/dev/null | sort -h | tail -15
sudo du -sh /var/backups/* /var/lib/shouon/storage /var/lib/postgresql 2>/dev/null

# CPU pressure: load average vs. core count
nproc && uptime

# database size per database
sudo -u postgres psql -c "SELECT datname, pg_size_pretty(pg_database_size(datname)) FROM pg_database ORDER BY pg_database_size(datname) DESC;"

Read the result like this:

Measurement Meaning
available consistently under ~200 MB genuinely out of RAM → upgrade or add swap (§14.6)
any oom-kill lines urgent — the kernel is killing processes, usually PostgreSQL
disk over ~80% reclaim first (§14.5); upgrade only if it is real data
load average > nproc for sustained periods CPU-bound → more cores
all comfortable do not upgrade. Save the money

Baseline for comparison (measured): 1.9 GB RAM with ~500 MB in use, ~34 GB disk at ~13%, Shouon production ≈ 470 MB. A second environment is already running on this box, which is where much of the remaining headroom went.

14.2 Ask RackNerd these exact questions, before paying

Open a ticket in the RackNerd client area (Support → Open Ticket) and ask:

  1. "I want to upgrade VPS from to . Is this done in place on the same node, or does it require a migration to a different node?"
  2. "Will my IPv4 address change?" ← the question that matters most
  3. "What is the expected downtime, and can I schedule the window?"
  4. "Is the disk expanded automatically, or do I need to grow the partition and filesystem myself afterwards?"
  5. "Is the upgrade billed pro-rata against my remaining term?"
🔴 Why question 2 decides everything. A new IP breaks three things on this server at once:
Breaks Why Fix
DOMAIN / DOMAIN_REDIRECT the A records point at the old IP new A records at the registrar, then §8.3
DOMAIN_SERVE=204.44.93.211.sslip.io the hostname contains the IP — the name itself changes set DOMAIN_SERVE=.sslip.io and update
any external IP allowlist — update at the other end
✅ And one thing that survives a new IP: the zrok tunnel. It is an outbound connection from our server to zrok's edge — no inbound port, no DNS of ours. So the zrok link keeps working through an IP change. That makes it the safe fallback to hand clients during the window.

14.3 Typical upgrade paths (confirm with support — see the warning above)

Path Downtime IP Notes
In-place RAM/CPU on the same node one reboot usually unchanged simplest; disk may still need growing (§14.7)
Migration to a different node longer, scheduled often changes treat as §14.4 in full
Build a new VPS and move you control it new IP the only path that keeps the old server as a rollback
🔑 If the budget allows, the "new VPS + migrate" path is the safest, because the old server stays untouched and serving until you cut over. It is also the only path where a failed upgrade costs nothing but time. Steps: install from scratch (deploy/install.sh), restore the newest backup, copy /var/lib/shouon/storage, copy deploy/.env, verify, then move DNS last.

14.4 The upgrade procedure, with a rollback at every step

# ── BEFORE ──────────────────────────────────────────────────────────────
# 1) [READ-ONLY] record the current state so you can compare afterwards
cd /opt/shouon && ./deploy/shouonctl env
./deploy/shouonctl check  2>&1 | tee /root/pre-upgrade-check.txt
free -h; df -h /; nproc; ip -4 addr show | grep inet

# 2) [WRITES] fresh backups of EVERY environment
cd /opt/shouon         && sudo ./deploy/shouonctl backup
cd /opt/shouon-staging && sudo ./deploy/shouonctl backup

# 3) [WRITES] copy the secrets and the backups OFF this server (from YOUR computer)
#    scp root@204.44.93.211:/opt/shouon/deploy/.env            ./shouon-env-backup
#    scp root@204.44.93.211:/var/backups/shouon/db-*.gz        ./
#    scp -r root@204.44.93.211:/var/lib/shouon/storage         ./storage-backup
#    → §11.4. Do this even for an in-place upgrade.

# 4) Take a provider snapshot if your plan offers one (RackNerd panel → Snapshots).
#    A snapshot is the only true one-click rollback.

# 5) Lower DNS TTL to 300 at the registrar, 24h in advance, if the IP may change.

# 6) [WRITES] stop writers cleanly so nothing is mid-transaction at reboot
cd /opt/shouon         && sudo ./deploy/shouonctl stop
cd /opt/shouon-staging && sudo ./deploy/shouonctl stop
sudo systemctl stop postgresql

# ── THE UPGRADE ─────────────────────────────────────────────────────────
# Performed by RackNerd, or by you in the panel. Wait for their confirmation.

# ── AFTER ───────────────────────────────────────────────────────────────
# 7) [READ-ONLY] did the hardware actually change? did the IP?
free -h; df -h /; nproc
ip -4 addr show | grep inet
curl -4 -s https://ifconfig.me; echo

# 8) [WRITES] bring it up (PostgreSQL first, always)
sudo systemctl start postgresql
cd /opt/shouon         && sudo ./deploy/shouonctl start
cd /opt/shouon-staging && sudo ./deploy/shouonctl start

# 9) [READ-ONLY] compare against the "before" file
cd /opt/shouon && ./deploy/shouonctl check 2>&1 | tee /root/post-upgrade-check.txt
diff /root/pre-upgrade-check.txt /root/post-upgrade-check.txt

Rollback: restore the provider snapshot (step 4), or — if you took the "new VPS" path — simply point DNS back at the old server, which never stopped running.

14.5 Before you pay: reclaim disk instead

[READ-ONLY] find it first:

sudo du -xh --max-depth=1 / 2>/dev/null | sort -h | tail -15
sudo du -sh /var/backups/* /var/log /var/cache/apt /var/lib/postgresql 2>/dev/null
sudo journalctl --disk-usage

[WRITES] safe reclaims, in increasing order of effect:

sudo apt-get clean                         # cached .deb packages
sudo journalctl --vacuum-time=14d          # old system logs
sudo apt-get autoremove --purge            # orphaned packages (read the list first)

[WRITES] Old backups — KEEP_DAYS in deploy/.env controls retention. Check what you are about to delete before deleting it:

ls -lt /var/backups/shouon | tail -20
⚠️ Leftover tables named _bk_* also sit in the production database (one is a copy of the permissions table). Deleting them is a data decision, not a cleanup — the owner decides.

14.6 🔴 There is no swap on this server — and that is a measured gap

[READ-ONLY] confirm:

free -h | grep -i swap
swapon --show

Measured: zero swap is configured anywhere, and nothing in the deployment scripts creates any. On a 1.9 GB box running PostgreSQL plus three Node services plus Caddy plus PHP, that means a memory spike does not slow down — the kernel kills a process outright, usually PostgreSQL.

A small swap file is far cheaper than a plan upgrade and often removes the need for one. [WRITES]:

sudo fallocate -l 2G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

# prefer RAM, use swap only under real pressure
echo 'vm.swappiness=10' | sudo tee /etc/sysctl.d/99-swap.conf
sudo sysctl --system

# [READ-ONLY] verify
free -h && swapon --show
⚠️ Its honest cost: swap lives on disk, so it consumes 2 GB of your ~34 GB, and a process that swaps heavily gets slow instead of being killed. For a database server, "slow but alive" is the right trade — but swap is a safety net, not capacity. If available is routinely near zero, you need real RAM. ✅ Do this before upgrading, then re-measure for a week (§14.1). If the oom-kill lines stop and available stays healthy, you have saved the upgrade.

14.7 After a real RAM or disk increase: retune, or the upgrade does nothing

New RAM is not used by PostgreSQL until you tell it. The installer set these for a 2 GB box:

Setting Current value Why
max_connections 50 —
shared_buffers 384 MB ≈ 20% of RAM
effective_cache_size 768 MB ≈ 40% of RAM (a hint to the planner)
work_mem 4 MB per sort, not per connection
maintenance_work_mem 96 MB —

[WRITES] On a 4 GB plan, for example:

sudo -u postgres psql -c "ALTER SYSTEM SET shared_buffers = '1GB';"
sudo -u postgres psql -c "ALTER SYSTEM SET effective_cache_size = '2GB';"
sudo -u postgres psql -c "ALTER SYSTEM SET maintenance_work_mem = '256MB';"
sudo systemctl restart postgresql          # shared_buffers needs a full restart

[READ-ONLY] verify they took effect:

sudo -u postgres psql -c "SHOW shared_buffers; SHOW effective_cache_size; SHOW work_mem;"
cd /opt/shouon && ./deploy/shouonctl check
🔴 Raise work_mem last and least. It is allocated per sort operation: 50 connections × 3 sorts × 16 MB is 2.4 GB of potential allocation. It is the best-known way to crash a small server, and it does not get safe just because you added RAM — it gets less obviously unsafe.

Disk: a bigger disk is often not visible until the partition and filesystem are grown. [READ-ONLY] check first:

lsblk && df -h /

If lsblk shows a device larger than the mounted filesystem, [WRITES]:

sudo growpart /dev/vda 1          # confirm the device/partition names from lsblk
sudo resize2fs /dev/vda1          # ext4;  use xfs_growfs / for XFS
df -h /                           # [READ-ONLY] verify
🔴 Read lsblk and confirm the names before typing them. Running growpart on the wrong device is destructive. If RackNerd said they expand it automatically, verify with df -h and do nothing.

14.8 A second server instead of a bigger one

If the reason to upgrade is a new project rather than Shouon's own growth, consider a second cheap VPS instead. The whole stack is designed for it: one project per server means no shared Caddy config, no shared PostgreSQL cluster, no port registry to maintain, and an outage in one project cannot take the other down. The cost comparison is usually close, and the blast radius is much smaller.


§15 — If you are Claude (or any AI agent) reading this

Treat the following as invariants. They were each learned from a production incident.

Before acting

  1. Read the source of any command before dictating it. Eight separate production rounds were lost to commands written from assumption. Read the script, and read the shell too (! expands inside double quotes in an interactive bash).
  2. A command is measured against the state of the server, not the repository. A file merged to main is not on the server until update pulls it; a file on a branch is not on main at all. hash → hash unchanged in update's output is how you detect this.
  3. State whether a command is read-only or writing, what it touches, what it does not, and whether it is reversible. The owner asks this every time and is right to.
  4. Ask which environment first. ./deploy/shouonctl env is read-only. user add does not ask, and once wrote an admin account into production from the staging copy.

While acting

  1. Always send update and check as a pair. Owner's standing instruction.
  2. update never runs SQL. Adding a block to an init/ file is half the work; running the file on the server is the other half and must be said explicitly. The source of truth for what to run is update's own trailing warning line.
  3. Deploy before SQL. The file does not exist on the server until then.
  4. One command at a time on anything destructive, with the output read back before the next.

When reporting

  1. Never claim success you did not measure. "Should work" is not a result.
  2. Distinguish three states, never two: measured-good / measured-bad / not measured. check prints ! for the third and says "and this is not a pass." Do not collapse it into either neighbour.
  3. Distinguish "the mechanism was open" from "harm occurred." Saying the first as if it were the second destroys the report's credibility; the owner has caught this.
  4. A correction written over a correct caution is worse than a wrong claim — it closes the file. If you reverse an earlier warning, quote the measurement that reversed it.
  5. Say the limits with the delivery, not after the complaint. "Measured in the lab" is not "confirmed in production," and "the owner said it works" is not "a guard measures it on production with a control."

When something fails

  1. When a guard fails, the first hypothesis is that the guard is wrong — this has held over thirty times in this project, most often because the fixture, not the condition, never passes through the failing state.
  2. Two checks that contradict each other mean one of them is wrong. The behavioural one (the one that measures the effect) is the honest one.
  3. If a revert "passes," inspect the revert and the measuring method before accusing the guard. Print how many occurrences the revert actually replaced; a revert that replaced zero did not happen.
  4. Read any line you generated programmatically after generating it. Nine incidents came from a generated line nobody read back.

Project rules

  1. Branch, PR, squash-merge to main immediately. Push only to the fork (Mahmoud-walid/shouon-al-ghithaa) — never the upstream repository. All commit and PR text in Arabic, explaining why, not what.
  2. Document every trap the moment you hit it, in the right file (CLAUDE.md, deploy/MULTI-PROJECT.md, deploy/test/README.md) — not only in the conversation.
  3. Correct yourself explicitly when you discover you were wrong.

§16 — Appendix

16.1 Quick reference card

# ── READ-ONLY: safe any time ─────────────────────────────────────────────
cd /opt/shouon && ./deploy/shouonctl env       # which environment am I in?
                  ./deploy/shouonctl status    # services
                  ./deploy/shouonctl check     # the full audit
                  ./deploy/shouonctl tunnel    # where does zrok point?
                  ./deploy/shouonctl user list
ss -tlnp | awk '{print $4}' | grep -oE '[0-9]+
  
  




 | sort -un   # ports in use
free -h && df -h /                                            # resources
caddy validate --config /etc/caddy/Caddyfile --adapter caddyfile
journalctl -u shouon-auth --since "-15 min" --no-pager | grep -a "x"

# ── WRITES ───────────────────────────────────────────────────────────────
sudo ./deploy/shouonctl update && sudo ./deploy/shouonctl check   # the deploy
sudo ./deploy/shouonctl backup
sudo ./deploy/shouonctl restart
sudo ./deploy/shouonctl tunnel here

# ── DESTRUCTIVE: back up and read the section first ──────────────────────
sudo ./deploy/shouonctl restore 
sudo ./deploy/shouonctl bootstrap
sudo -u postgres dropdb 

16.2 File and directory map

Path What In git?
/opt/shouon production code yes (a clone)
/opt/shouon-staging staging code yes (a separate clone)
/opt//deploy/.env secrets — JWT, anon key, vault key, DB passwords no
/etc//postgrest.conf DB URI + JWT for the REST API no
/etc//auth.env DB URI + JWT + vault key no
/etc//storage.env JWT + storage root + upload limit no
/var/lib//storage/ uploaded files no
/var/backups// database backups no
/etc/caddy/Caddyfile server-wide — never edit generated
/etc/caddy/conf.d/.caddy one project's web config generated
/var/lib/caddy/ issued TLS certificates — never delete no
/etc/cron.d/ that environment's scheduled jobs generated
/etc/cron.d/shouon-host shared machine disk monitor generated
/opt/openziti/etc/zrok/*.env zrok share config (target port, unique name) no
/var/lib/private/zrok-share/.zrok/ real zrok state (DynamicUser) no

16.3 Where to read more (in the repository)

File When
deploy/README.md installation from scratch, command reference, troubleshooting
deploy/MULTI-PROJECT.md the authoritative source for everything in §5–§8
deploy/test/README.md the test guards and their known failure modes
security/README.md the order the hardening SQL files must be applied in
deploy/shouonctl every command and exactly what update does
deploy/scripts/env.sh how the environment is resolved — read this once
deploy/scripts/new-env.sh the staging creation script and its preflight
deploy/scripts/render-caddy.sh how domains become Caddy site blocks
CLAUDE.md the full incident log — every trap, with its measurement

Written from the server's own scripts and configuration, not from memory. Every number marked "measured" came from the server. Where something was not measured, it says so.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论