mirror of
https://github.com/netbirdio/docs.git
synced 2026-09-28 17:59:05 +02:00
Restructure Troubleshooting into a hub with per-area pages (#814)
* Restructure Troubleshooting into a hub with per-area pages - Add a Troubleshooting hub (/help/troubleshooting) with icon/chip cards and a "Still stuck?" CTA - Split NetBird Client troubleshooting into an overview + per-OS pages (Linux, Windows, macOS, Android, iOS) - Split Self-hosted troubleshooting into an overview + per-area pages (installation, IdP, dashboard, certificates, connectivity, database) - Split "Report bugs and issues" into Community Support and NetBird Support pages - Add Troubleshooting resource connectivity and a NetBird Cloud pending-approval page - Add DNS troubleshooting Issue 8 (Windows NRPT rule blocked by a lingering GPO) - Cross-reference the new pages from networks, DNS, and reverse-proxy docs; update nav Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Address review: client terminology, dead props, labels, cross-links - Use "client" instead of "agent" across the client troubleshooting pages (headings, prose, anchors) - Remove unused source: props from the Troubleshooting hub tiles - Relabel the "NetBird Cloud" grouping to "Cloud & identity" (SSO/provisioning also apply to self-hosted) - Add a Tiles title on the report-bug landing; add reverse-proxy -> resource-connectivity cross-link - Fix comma splices introduced by the em-dash cleanup in relayed-connections Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Add client-side hash redirect for moved self-hosted anchors Old deep links like /selfhosted/troubleshooting#debugging-turn-connections now forward to the per-area page, since next.config redirects can't act on the URL fragment. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Apply docs-skill review: conventions + reshape area pages - "open source" (no hyphen), expand NRPT on first use, descriptive alt text + captions on TURN images - Fix inherited "Netbird" casing in the client glossary - Reshape the six self-hosted area pages to Symptom -> likely causes (ordered) -> Fix -> Confirm, preserving anchored headings Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Fix two typos in client glossary (CodeRabbit) - "nunning" -> "running" in the glossary - possessive "it's" -> "its" in the routing-table sentence Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: fix two broken links in troubleshooting pages - database: point the "upgrade path" link at /selfhosted/maintenance/upgrade; selfhosted-quickstart has no #upgrade anchor so the old link landed at page top - client: add HashRedirect so old #net-bird-agent-status deep links forward to the renamed #net-bird-client-status section on the same page * docs: address review follow-ups (deep-link redirects + client casing) - self-hosted troubleshooting: extend the HashRedirect map with the per-issue (###-level) anchors from the old single page, so old deep links land on the exact sub-section of the new area page rather than just the page top - client glossary: lowercase "NetBird client" in the peer-a/peer-b entries (house convention) and fix "linux" -> "Linux" * docs: review polish — fix image class + first-use acronym glosses - connectivity: fix bad CSS class imagewrapper-nig -> imagewrapper on the TURN-test screenshot (the typo'd class matched no style and broke zoom) - gloss acronyms on first use: GPO (DNS Issue 8), IdP/SSO (identity-provider), ACME (certificates), CORS (dashboard) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Jack Carter <128555021+SunsetDrifter@users.noreply.github.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
Jack Carter
parent
ffb72cfa34
commit
5729ad035e
@@ -68,7 +68,7 @@ After you answer the prompts, the script writes a `docker-compose.yml` with the
|
||||
|
||||
The readiness check probes `https://your-domain/oauth2/.well-known/openid-configuration` through the proxy. The script auto-detects your Traefik container (any container running a Traefik image with ports 80 and 443 published) so it can pull diagnostic logs if something goes wrong.
|
||||
|
||||
If the wait check hangs, see [Installation Script Issues](/selfhosted/troubleshooting#installation-script-issues) for the common causes.
|
||||
If the wait check hangs, see [Installation script issues](/selfhosted/troubleshooting/installation) for the common causes.
|
||||
|
||||
### Existing Deployments
|
||||
|
||||
|
||||
@@ -1,255 +1,102 @@
|
||||
import { TroubleshootingStart } from "@/components/TroubleshootingStart"
|
||||
import { TroubleshootingTiles } from "@/components/TroubleshootingTiles"
|
||||
import { HashRedirect } from "@/components/HashRedirect"
|
||||
|
||||
export const description =
|
||||
"Troubleshoot a self-hosted NetBird deployment: collect diagnostics, then jump to the area that matches your issue."
|
||||
|
||||
<HashRedirect
|
||||
map={{
|
||||
"embedded-idp-issues": "/selfhosted/troubleshooting/identity-provider",
|
||||
"setup-page-not-accessible": "/selfhosted/troubleshooting/identity-provider#setup-page-not-accessible",
|
||||
"setup-already-completed-error-http-412": "/selfhosted/troubleshooting/identity-provider#setup-already-completed-error-http-412",
|
||||
"password-not-working-after-user-creation": "/selfhosted/troubleshooting/identity-provider#password-not-working-after-user-creation",
|
||||
"sso-connector-not-appearing-on-login-page": "/selfhosted/troubleshooting/identity-provider#sso-connector-not-appearing-on-login-page",
|
||||
"invalid-redirect-uri-error-from-id-p": "/selfhosted/troubleshooting/identity-provider#invalid-redirect-uri-error-from-id-p",
|
||||
"identity-providers-tab-not-visible": "/selfhosted/troubleshooting/identity-provider#identity-providers-tab-not-visible",
|
||||
"users-not-syncing-from-sso-provider": "/selfhosted/troubleshooting/identity-provider#users-not-syncing-from-sso-provider",
|
||||
"installation-script-issues": "/selfhosted/troubleshooting/installation",
|
||||
"script-hangs-on-waiting-for-net-bird-server-to-become-ready": "/selfhosted/troubleshooting/installation#script-hangs-on-waiting-for-net-bird-server-to-become-ready",
|
||||
"script-fails-on-existing-installation-check": "/selfhosted/troubleshooting/installation#script-fails-on-existing-installation-check",
|
||||
"debugging-turn-connections": "/selfhosted/troubleshooting/connectivity#debugging-turn-connections",
|
||||
"connection-issues": "/selfhosted/troubleshooting/connectivity",
|
||||
"peers-cant-connect-to-each-other": "/selfhosted/troubleshooting/connectivity#peers-cant-connect-to-each-other",
|
||||
"management-service-unreachable": "/selfhosted/troubleshooting/connectivity#management-service-unreachable",
|
||||
"dashboard-issues": "/selfhosted/troubleshooting/dashboard",
|
||||
"dashboard-shows-blank-page": "/selfhosted/troubleshooting/dashboard#dashboard-shows-blank-page",
|
||||
"unauthorized-or-403-errors": "/selfhosted/troubleshooting/dashboard#unauthorized-or-403-errors",
|
||||
"certificate-issues": "/selfhosted/troubleshooting/certificates",
|
||||
"lets-encrypt-certificate-not-renewing": "/selfhosted/troubleshooting/certificates#lets-encrypt-certificate-not-renewing",
|
||||
"certificate-errors-with-custom-reverse-proxy": "/selfhosted/troubleshooting/certificates#certificate-errors-with-custom-reverse-proxy",
|
||||
"database-issues": "/selfhosted/troubleshooting/database",
|
||||
"management-service-wont-start-after-upgrade": "/selfhosted/troubleshooting/database#management-service-wont-start-after-upgrade",
|
||||
"data-corruption-after-power-loss": "/selfhosted/troubleshooting/database#data-corruption-after-power-loss",
|
||||
}}
|
||||
/>
|
||||
|
||||
# Troubleshooting
|
||||
|
||||
This page will help with various issues when self-hosting NetBird.
|
||||
|
||||
## Embedded IdP Issues
|
||||
|
||||
### Setup page not accessible
|
||||
|
||||
**Problem**: You can't access the `/setup` page to create the first user.
|
||||
|
||||
**Solutions**:
|
||||
- The `/setup` page is only available when no users exist. If you've already created a user, use the regular login page.
|
||||
- Check that the embedded IdP is enabled in your configuration.
|
||||
- Verify the Management service is running: `docker compose logs management`
|
||||
|
||||
### "Setup already completed" error (HTTP 412)
|
||||
|
||||
**Problem**: The setup endpoint returns a 412 error.
|
||||
|
||||
**Solution**: Setup has already been completed. The first user was already created. Use the regular login page to sign in.
|
||||
|
||||
### Password not working after user creation
|
||||
|
||||
**Problem**: You created a user but the password doesn't work.
|
||||
|
||||
**Solutions**:
|
||||
- Passwords are shown only once during user creation. If you didn't save it, you'll need to delete and recreate the user.
|
||||
- Via API, you can create a new user with a new password.
|
||||
- For the owner account, you may need to reset the database if no other admins exist.
|
||||
|
||||
### SSO connector not appearing on login page
|
||||
|
||||
**Problem**: You configured an identity provider but it doesn't show on the login page.
|
||||
|
||||
**Solutions**:
|
||||
1. Verify the connector was saved: Go to **Settings** → **Identity Providers**
|
||||
2. Check that the redirect URL is correctly configured in your IdP
|
||||
3. Review Management service logs for configuration errors: `docker compose logs management`
|
||||
4. Ensure the IdP application has the correct redirect URI from NetBird
|
||||
|
||||
### "Invalid redirect URI" error from IdP
|
||||
|
||||
**Problem**: When clicking an SSO button, the IdP returns an invalid redirect URI error.
|
||||
|
||||
**Solutions**:
|
||||
1. Copy the exact redirect URL from NetBird (shown after saving the connector)
|
||||
2. Add it to your IdP's allowed redirect URIs
|
||||
3. Check for trailing slashes or typos
|
||||
4. Some IdPs are case-sensitive
|
||||
|
||||
### Identity Providers tab not visible
|
||||
|
||||
**Problem**: You don't see the Identity Providers tab in Settings.
|
||||
|
||||
**Solution**: This tab is only visible when the embedded IdP is enabled. Check your deployment configuration:
|
||||
- For quickstart deployments, the embedded IdP should be enabled by default
|
||||
- For the combined setup, the embedded IdP is always enabled
|
||||
- For older multi-container deployments, ensure `EmbeddedIdP.Enabled` is set to `true` in `management.json`
|
||||
|
||||
### Users not syncing from SSO provider
|
||||
|
||||
**Problem**: Users who authenticate via SSO don't appear in the user list.
|
||||
|
||||
**Solutions**:
|
||||
- Users appear after their first successful login, not immediately after connector configuration
|
||||
- Check that the SSO flow completes successfully (user should land on Dashboard)
|
||||
- Review Management logs for any token validation errors
|
||||
|
||||
## Installation Script Issues
|
||||
|
||||
### Script hangs on "Waiting for NetBird server to become ready"
|
||||
|
||||
**Problem**: The `getting-started.sh` script stays on the wait line for several minutes, even though netbird-server itself looks healthy in `docker ps`.
|
||||
|
||||
The wait check probes the OIDC endpoint through your reverse proxy, so a healthy server alone isn't enough. Something on the path from the public internet to netbird-server is broken. Check these in order:
|
||||
|
||||
**1. DNS**
|
||||
|
||||
Confirm your domain points at the host:
|
||||
|
||||
```bash
|
||||
dig +short netbird.example.com
|
||||
```
|
||||
|
||||
The result should be the public IP of the VM running NetBird. If it's empty or wrong, the cert challenge can't complete and the probe will never succeed.
|
||||
|
||||
**2. Cert issuance**
|
||||
|
||||
For options 0 (built-in Traefik) and 1 (existing Traefik), check that a real cert was issued:
|
||||
|
||||
```bash
|
||||
curl -vI https://netbird.example.com 2>&1 | grep -E "subject|issuer"
|
||||
```
|
||||
|
||||
If you see a self-signed or default cert, ACME hasn't completed. Check Traefik logs for ACME errors. Common causes: port 80 not reachable from the internet, DNS not propagated yet, or Let's Encrypt rate limit hit.
|
||||
|
||||
**3. Traefik can talk to the Docker socket (option 1 only)**
|
||||
|
||||
Traefik discovers NetBird containers via Docker labels. If the Docker socket is mounted but Traefik can't read it (old SDK, wrong path, permission denied), routes never get created. Check Traefik logs for errors like `client version is too old` or `permission denied while trying to connect to the Docker daemon socket`.
|
||||
|
||||
**4. Network membership matches (option 1 only)**
|
||||
|
||||
Confirm the Traefik container and NetBird containers are actually on the same Docker network:
|
||||
|
||||
```bash
|
||||
docker network inspect <your-network-name> | grep -A 2 Containers
|
||||
```
|
||||
|
||||
You should see both Traefik and the NetBird containers listed.
|
||||
|
||||
**5. Routing works end to end**
|
||||
|
||||
If the four above look fine, test the full path manually:
|
||||
|
||||
```bash
|
||||
curl -k https://netbird.example.com/oauth2/.well-known/openid-configuration
|
||||
```
|
||||
|
||||
A valid response is JSON containing `"issuer"`. Anything else (empty body, 404, 502, connection refused) tells you where to dig further: 502 means Traefik can't reach netbird-server, 404 means the labels aren't matched, connection refused means you're not hitting Traefik at all.
|
||||
|
||||
You can let the script keep waiting while you debug. As soon as the probe succeeds, the script will continue. If you'd rather start over, kill it with Ctrl+C, run `docker compose down -v`, fix the issue, and re-run.
|
||||
|
||||
### Script fails on existing installation check
|
||||
|
||||
**Problem**: The script exits immediately with a message about generated files already existing.
|
||||
|
||||
**Solution**: This is intentional and protects you from overwriting a working setup. To start fresh:
|
||||
|
||||
```bash
|
||||
docker compose down --volumes
|
||||
rm -f docker-compose.yml dashboard.env config.yaml proxy.env \
|
||||
traefik-dynamic.yaml nginx-netbird.conf caddyfile-netbird.txt \
|
||||
npm-advanced-config.txt
|
||||
```
|
||||
|
||||
Then re-run the script. Be aware this removes all NetBird data including users and peers.
|
||||
|
||||
## Debugging TURN connections
|
||||
|
||||
In the case that the peer-to-peer connection is not an option then the peer will use the TURN server for the secure connection establishment. If the connection is not possible even with TURN (Relay),
|
||||
then we need to confirm that your turn configuration is correct and that it is available.
|
||||
|
||||
To test your TURN configuration you can access the [online tester](https://webrtc.github.io/samples/src/content/peerconnection/trickle-ice).
|
||||
There you will find a ICE servers input box, where you can select and remove the existing server, then add your turn server
|
||||
configuration as follows:
|
||||
|
||||
Please replace <b>netbird.DOMAIN.com</b> and <b>PASSWORD</b> with your STUN/TURN server details from your configuration (<b>config.yaml</b> for the combined setup, or the <b>TURNConfig</b> section in <b>management.json</b> for older multi-container setups), then click on <b>Add server</b>.
|
||||
|
||||
<p>
|
||||
<img src="/docs-static/img/selfhosted/troubleshooting/turn.png" alt="turn" width="700" className="imagewrapper"/>
|
||||
</p>
|
||||
|
||||
You should see an output similar to the following:
|
||||
<p>
|
||||
<img src="/docs-static/img/selfhosted/troubleshooting/turn-test-out.png" alt="turn" width="700" className="imagewrapper-nig"/>
|
||||
</p>
|
||||
Where you have the following types: `host` (local address), `srflx` (STUN reflexive address), `relay`
|
||||
(TURN relay address). If `srflx` and `relay` are not present then the TURN server is not working or not accessible and you should review the required ports in the [requirements section](/selfhosted/selfhosted-guide#requirements).
|
||||
|
||||
## Dashboard Issues
|
||||
|
||||
### Dashboard shows blank page
|
||||
|
||||
**Problem**: The Dashboard loads but shows a blank page or errors.
|
||||
|
||||
**Solutions**:
|
||||
1. Check browser console for JavaScript errors (F12 → Console)
|
||||
2. Verify the Dashboard can reach the Management API
|
||||
3. Check CORS configuration if running behind a custom reverse proxy
|
||||
4. Clear browser cache and try again
|
||||
|
||||
### "Unauthorized" or "403" errors
|
||||
|
||||
**Problem**: API calls return unauthorized or forbidden errors.
|
||||
|
||||
**Solutions**:
|
||||
1. Verify your authentication token is valid
|
||||
2. Check that the user has appropriate permissions
|
||||
3. For API access, ensure you're using a valid Personal Access Token (PAT)
|
||||
4. Review Management service logs for detailed error information
|
||||
|
||||
## Certificate Issues
|
||||
|
||||
### Let's Encrypt certificate not renewing
|
||||
|
||||
**Problem**: SSL certificate expires and doesn't auto-renew.
|
||||
|
||||
**Solutions**:
|
||||
1. Ensure port 80 is accessible for ACME challenge
|
||||
2. Check Caddy logs: `docker compose logs caddy`
|
||||
3. Verify the domain points to the correct IP
|
||||
4. Manually trigger renewal: `docker exec -it netbird-caddy caddy reload`
|
||||
|
||||
### Certificate errors with custom reverse proxy
|
||||
|
||||
**Problem**: SSL errors when using your own reverse proxy.
|
||||
|
||||
**Solutions**:
|
||||
1. Ensure your reverse proxy terminates SSL correctly
|
||||
2. Set `NETBIRD_DISABLE_LETSENCRYPT=true` in your configuration
|
||||
3. Configure proper headers (X-Forwarded-For, X-Forwarded-Proto)
|
||||
4. Verify HTTP/2 support is enabled for gRPC endpoints
|
||||
|
||||
## Connection Issues
|
||||
|
||||
### Peers can't connect to each other
|
||||
|
||||
**Problem**: Peers are registered but can't establish connections.
|
||||
|
||||
**Solutions**:
|
||||
1. Check that UDP port 3478 is accessible (STUN/TURN)
|
||||
2. Verify the TURN server is working (see [TURN debugging](#debugging-turn-connections))
|
||||
3. Check firewall rules on both peers
|
||||
4. Review peer logs: `netbird status -d`
|
||||
|
||||
### Management service unreachable
|
||||
|
||||
**Problem**: Clients can't connect to the Management service.
|
||||
|
||||
**Solutions**:
|
||||
1. Verify port 443 is accessible
|
||||
2. Check DNS resolution for your domain
|
||||
3. Review Management logs: `docker compose logs management`
|
||||
4. Test with curl: `curl -v https://your-domain.com/api/health`
|
||||
|
||||
## Database Issues
|
||||
|
||||
### Management service won't start after upgrade
|
||||
|
||||
**Problem**: After upgrading, the Management service fails to start.
|
||||
|
||||
**Solutions**:
|
||||
1. Check logs for migration errors: `docker compose logs management`
|
||||
2. Ensure you followed the [upgrade path](/selfhosted/selfhosted-quickstart#upgrade) for your version
|
||||
3. Restore from backup if needed
|
||||
4. For major version jumps, you may need intermediate upgrades
|
||||
|
||||
### Data corruption after power loss
|
||||
|
||||
**Problem**: Services don't start properly after unexpected shutdown.
|
||||
|
||||
**Solutions**:
|
||||
1. Check for database lock files
|
||||
2. Review all service logs
|
||||
3. Consider restoring from backup
|
||||
4. For SQLite databases, you may need to run integrity checks
|
||||
This page helps with issues when self-hosting NetBird. Collect the diagnostics below, then pick the area that matches your problem.
|
||||
|
||||
<TroubleshootingStart
|
||||
eyebrow="Start here"
|
||||
title="Collect diagnostics first"
|
||||
description="Almost every self-hosted issue is faster to resolve with service status and logs in hand. Grab these before you go deeper."
|
||||
steps={[
|
||||
{ label: "1. Check services", title: "Service status", command: "docker compose ps", hint: "Are management, signal, relay, and the dashboard all Up?" },
|
||||
{ label: "2. Read the logs", title: "Service logs", command: "docker compose logs -f management", hint: "Most config, auth, and migration errors surface here." },
|
||||
{ label: "3. Test reachability", title: "Proxy to API", command: "curl -v https://YOUR_DOMAIN/api/health", hint: "Confirms the reverse proxy can reach the Management API." },
|
||||
]}
|
||||
/>
|
||||
|
||||
<TroubleshootingTiles
|
||||
title="Find your issue by area"
|
||||
id="find-your-issue-by-area"
|
||||
description="Each area lives on its own page. Pick the one closest to your problem."
|
||||
items={[
|
||||
{
|
||||
title: "Installation",
|
||||
href: "/selfhosted/troubleshooting/installation",
|
||||
icon: "terminal",
|
||||
description: "The getting-started script: the readiness wait, DNS, certificates, Traefik, and re-running cleanly.",
|
||||
},
|
||||
{
|
||||
title: "Embedded IdP",
|
||||
href: "/selfhosted/troubleshooting/identity-provider",
|
||||
icon: "shield",
|
||||
description: "Setup page access, SSO connectors, redirect URIs, and users syncing from your provider.",
|
||||
},
|
||||
{
|
||||
title: "Dashboard",
|
||||
href: "/selfhosted/troubleshooting/dashboard",
|
||||
icon: "globe",
|
||||
description: "Blank pages and unauthorized or 403 errors in the self-hosted dashboard.",
|
||||
},
|
||||
{
|
||||
title: "Certificates",
|
||||
href: "/selfhosted/troubleshooting/certificates",
|
||||
icon: "lock",
|
||||
description: "Let's Encrypt renewal and TLS errors behind a custom reverse proxy.",
|
||||
},
|
||||
{
|
||||
title: "Connectivity",
|
||||
href: "/selfhosted/troubleshooting/connectivity",
|
||||
icon: "firewall",
|
||||
description: "Testing the TURN server, peer-to-peer failures, and an unreachable Management service.",
|
||||
},
|
||||
{
|
||||
title: "Database",
|
||||
href: "/selfhosted/troubleshooting/database",
|
||||
icon: "database",
|
||||
description: "The Management service failing to start after an upgrade, and recovery after power loss.",
|
||||
},
|
||||
]}
|
||||
/>
|
||||
|
||||
## Getting Help
|
||||
|
||||
If you're still experiencing issues:
|
||||
If you're still experiencing issues, see [Report bugs and issues](/help/report-bug-issues) for the right channel:
|
||||
|
||||
1. **Check logs**: `docker compose logs` for all services
|
||||
2. **Search existing issues**: [GitHub Issues](https://github.com/netbirdio/netbird/issues)
|
||||
3. **Join our community**: [Slack Channel](/slack-url)
|
||||
4. **Open an issue**: Include logs, configuration (without secrets), and steps to reproduce
|
||||
1. **Gather evidence first**: `docker compose logs` for all services, your configuration (without secrets), and the steps to reproduce.
|
||||
2. **Open source self-hosted and general questions** go to [Community Support](/help/community-support): the [Slack Channel](/slack-url) for quick questions, or [GitHub Discussions](https://github.com/netbirdio/netbird/discussions) for a written record.
|
||||
3. **Cloud customers and users, and commercial-license deployments** can reach the team through [NetBird Support](/help/netbird-support).
|
||||
|
||||
@@ -0,0 +1,30 @@
|
||||
export const description =
|
||||
"Self-hosted NetBird certificate troubleshooting: Let's Encrypt renewal and TLS errors behind a custom reverse proxy."
|
||||
|
||||
# Certificate issues
|
||||
|
||||
TLS and certificate problems on a self-hosted deployment. For other areas, start from [Troubleshooting](/selfhosted/troubleshooting).
|
||||
|
||||
## Let's Encrypt certificate not renewing
|
||||
|
||||
**Symptom**: The TLS certificate expires and does not auto-renew, so clients and browsers report an expired or invalid certificate.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **Port 80 is not reachable from the internet.** The ACME HTTP challenge (how Let's Encrypt validates your domain) needs inbound TCP/80. Confirm your firewall and cloud security groups allow it.
|
||||
2. **The domain no longer points at this host.** Verify the `A`/`AAAA` record resolves to the server's public IP.
|
||||
3. **A renewal error in the proxy.** Check the certificate manager's logs: `docker compose logs caddy`. If needed, force a reload: `docker exec -it netbird-caddy caddy reload`.
|
||||
|
||||
**Confirm**: `curl -vI https://YOUR_DOMAIN 2>&1 | grep -E "issuer|expire"` shows a current Let's Encrypt certificate.
|
||||
|
||||
## Certificate errors with custom reverse proxy
|
||||
|
||||
**Symptom**: TLS errors when terminating TLS on your own reverse proxy instead of the bundled one.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **Let's Encrypt is still enabled, so two components fight over TLS.** Set `NETBIRD_DISABLE_LETSENCRYPT=true` so NetBird stops managing certificates and leaves termination to your proxy.
|
||||
2. **Forwarded headers are missing.** Set `X-Forwarded-For` and `X-Forwarded-Proto` on the proxy so NetBird sees the original scheme and client.
|
||||
3. **gRPC fails without HTTP/2.** The Management gRPC endpoints need HTTP/2; enable it on the proxy.
|
||||
|
||||
**Confirm**: The dashboard loads over your proxy without TLS warnings, and `netbird status` from a client shows `Management: Connected`.
|
||||
@@ -0,0 +1,52 @@
|
||||
export const description =
|
||||
"Self-hosted NetBird connectivity troubleshooting: testing the TURN server, peer-to-peer connection failures, and an unreachable Management service."
|
||||
|
||||
# Connectivity issues
|
||||
|
||||
Peer connectivity and relay problems on a self-hosted deployment. For other areas, start from [Troubleshooting](/selfhosted/troubleshooting).
|
||||
|
||||
## Debugging TURN connections
|
||||
|
||||
When peers can't establish a direct connection, they fall back to the TURN (relay) server. If the connection still fails with relay, confirm your TURN configuration is correct and reachable.
|
||||
|
||||
To test it, open the [Trickle ICE test tool](https://webrtc.github.io/samples/src/content/peerconnection/trickle-ice). In the **ICE servers** box, remove the default server and add your TURN server. Replace <b>netbird.DOMAIN.com</b> and <b>PASSWORD</b> with your STUN/TURN details from your configuration (`config.yaml` for the combined setup, or the `TURNConfig` section in `management.json` for older multi-container setups), then click <b>Add server</b>.
|
||||
|
||||
<p>
|
||||
<img src="/docs-static/img/selfhosted/troubleshooting/turn.png" alt="The Trickle ICE test tool with a NetBird STUN/TURN server added to the ICE servers list" width="700" className="imagewrapper"/>
|
||||
</p>
|
||||
|
||||
*Add your STUN/TURN server in the ICE servers box, then click Add server.*
|
||||
|
||||
Gather candidates and read the result:
|
||||
|
||||
<p>
|
||||
<img src="/docs-static/img/selfhosted/troubleshooting/turn-test-out.png" alt="Trickle ICE test output listing host, srflx, and relay candidate types" width="700" className="imagewrapper"/>
|
||||
</p>
|
||||
|
||||
*A working TURN server returns `srflx` and `relay` candidates.*
|
||||
|
||||
The candidate types are `host` (local address), `srflx` (STUN reflexive address), and `relay` (TURN relay address). If `srflx` and `relay` are missing, the TURN server is not working or not reachable. Review the required ports in the [requirements section](/selfhosted/selfhosted-guide#requirements).
|
||||
|
||||
## Peers can't connect to each other
|
||||
|
||||
**Symptom**: Peers are registered and show in the dashboard, but can't establish a connection between them.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **STUN/TURN is unreachable.** Confirm UDP/3478 (STUN/TURN) is open, then verify the relay path with the [TURN test above](#debugging-turn-connections).
|
||||
2. **A host firewall is dropping traffic** on one or both peers. Check the local firewall rules on each peer.
|
||||
3. **The peer itself is unhealthy.** Run `netbird status -d` on both peers and check the connection type and last handshake.
|
||||
|
||||
**Confirm**: `netbird status -d` on both peers shows `Status: Connected` and a recent WireGuard handshake.
|
||||
|
||||
## Management service unreachable
|
||||
|
||||
**Symptom**: Clients can't register or stay connected, reporting that the Management service is unreachable.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **TCP/443 is blocked** between the client and the server. Confirm the port is open end to end.
|
||||
2. **DNS doesn't resolve your domain** to the server's public IP. Check resolution from the client.
|
||||
3. **Management is down or erroring.** Review `docker compose logs management`, and test the endpoint directly: `curl -v https://YOUR_DOMAIN/api/health`.
|
||||
|
||||
**Confirm**: `curl https://YOUR_DOMAIN/api/health` returns a healthy response, and `netbird status` shows `Management: Connected`.
|
||||
@@ -0,0 +1,30 @@
|
||||
export const description =
|
||||
"Self-hosted NetBird dashboard troubleshooting: blank pages and unauthorized or 403 errors."
|
||||
|
||||
# Dashboard issues
|
||||
|
||||
Problems with the self-hosted dashboard. For other areas, start from [Troubleshooting](/selfhosted/troubleshooting).
|
||||
|
||||
## Dashboard shows blank page
|
||||
|
||||
**Symptom**: The dashboard loads but renders a blank page, sometimes with errors in the browser console.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **The dashboard can't reach the Management API.** Open the browser console (F12 → Console) and look for failed requests. Confirm the dashboard's configured API URL is correct and reachable from the browser.
|
||||
2. **CORS (Cross-Origin Resource Sharing) is blocking the API behind a custom reverse proxy.** Serve the API on the same origin as the dashboard, or set the correct CORS headers on your proxy.
|
||||
3. **A stale cached bundle.** Hard-reload the page or clear the browser cache.
|
||||
|
||||
**Confirm**: Reload the dashboard. The login or peers view renders, and the console shows no failed API requests.
|
||||
|
||||
## "Unauthorized" or "403" errors
|
||||
|
||||
**Symptom**: API calls return 401 Unauthorized or 403 Forbidden.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **An expired or invalid token.** Re-authenticate. For direct API access, use a valid Personal Access Token (PAT).
|
||||
2. **The user lacks permission** for the action. Check the user's role in the dashboard.
|
||||
3. **Management can't validate the token.** Review `docker compose logs management` for token-validation errors.
|
||||
|
||||
**Confirm**: Re-run the action. It succeeds, and the Management logs show no auth errors.
|
||||
@@ -0,0 +1,30 @@
|
||||
export const description =
|
||||
"Self-hosted NetBird database troubleshooting: the Management service failing to start after an upgrade, and recovery after power loss."
|
||||
|
||||
# Database issues
|
||||
|
||||
Database and migration problems on a self-hosted deployment. For other areas, start from [Troubleshooting](/selfhosted/troubleshooting).
|
||||
|
||||
## Management service won't start after upgrade
|
||||
|
||||
**Symptom**: After an upgrade, the Management container fails to start or crash-loops.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **A failed schema migration.** Check `docker compose logs management` for migration errors. This is the usual cause and the logs name the failing step.
|
||||
2. **A skipped intermediate version.** Major version jumps may require stepping through intermediate releases. Follow the [upgrade path](/selfhosted/maintenance/upgrade) for your version.
|
||||
3. **A corrupt or partially-migrated database.** If the migration can't be completed, restore from a pre-upgrade backup and retry the upgrade.
|
||||
|
||||
**Confirm**: `docker compose ps` shows `management` as `Up`, and its logs end with the service listening rather than a migration error.
|
||||
|
||||
## Data corruption after power loss
|
||||
|
||||
**Symptom**: Services don't start cleanly after an unexpected shutdown.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **A stale lock file.** Check for and remove a leftover database lock if the process that held it is gone.
|
||||
2. **An interrupted write.** Review all service logs to find which store is failing. For SQLite, run an integrity check on the database file.
|
||||
3. **An unrecoverable file.** If the store can't be repaired, restore from the most recent backup.
|
||||
|
||||
**Confirm**: All services report `Up` in `docker compose ps`, and peers reconnect with `netbird status` showing `Management: Connected`.
|
||||
@@ -0,0 +1,80 @@
|
||||
export const description =
|
||||
"Self-hosted NetBird embedded IdP and SSO troubleshooting: setup page access, redirect URIs, connectors, and user sync."
|
||||
|
||||
# Embedded IdP issues
|
||||
|
||||
Problems with the embedded identity provider (IdP) and SSO (single sign-on) on a self-hosted deployment. For other areas, start from [Troubleshooting](/selfhosted/troubleshooting).
|
||||
|
||||
## Setup page not accessible
|
||||
|
||||
**Symptom**: You can't open the `/setup` page to create the first user.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **A user already exists.** `/setup` is only available when no users exist. If you've already created one, use the regular login page.
|
||||
2. **The embedded IdP is disabled.** Confirm it is enabled in your configuration.
|
||||
3. **Management isn't running.** Check `docker compose logs management`.
|
||||
|
||||
**Confirm**: `/setup` loads (first run), or the regular login page loads (setup already done).
|
||||
|
||||
## "Setup already completed" error (HTTP 412)
|
||||
|
||||
**Symptom**: The setup endpoint returns a 412 error.
|
||||
|
||||
**Cause**: Setup has already completed; the first user exists.
|
||||
|
||||
**Fix**: Use the regular login page to sign in.
|
||||
|
||||
## Password not working after user creation
|
||||
|
||||
**Symptom**: You created a user but the password doesn't work.
|
||||
|
||||
**Cause**: Passwords are shown only once at creation. If it wasn't saved, it can't be recovered.
|
||||
|
||||
**Fix**: Recreate the user (delete and re-add, or create a new one with a new password via the API). For the owner account with no other admins, you may need to reset the database.
|
||||
|
||||
**Confirm**: You can log in with the new credentials.
|
||||
|
||||
## SSO connector not appearing on login page
|
||||
|
||||
**Symptom**: You configured an identity provider but it doesn't show on the login page.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **The connector wasn't saved.** Check **Settings → Identity Providers**.
|
||||
2. **The redirect URI is wrong.** Ensure the IdP application has the exact redirect URI shown by NetBird.
|
||||
3. **A configuration error.** Review `docker compose logs management` for connector errors.
|
||||
|
||||
**Confirm**: The SSO button appears on the login page and starts the IdP flow.
|
||||
|
||||
## "Invalid redirect URI" error from IdP
|
||||
|
||||
**Symptom**: Clicking an SSO button returns an invalid redirect URI error from the IdP.
|
||||
|
||||
**Cause**: The redirect URI registered at the IdP doesn't exactly match the one NetBird sends.
|
||||
|
||||
**Fix**: Copy the exact redirect URL from NetBird (shown after saving the connector) into your IdP's allowed redirect URIs. Watch for trailing slashes, typos, and case sensitivity.
|
||||
|
||||
**Confirm**: The SSO flow completes and lands you on the dashboard.
|
||||
|
||||
## Identity Providers tab not visible
|
||||
|
||||
**Symptom**: The Identity Providers tab is missing from Settings.
|
||||
|
||||
**Cause**: The tab only appears when the embedded IdP is enabled.
|
||||
|
||||
**Fix**: Check your deployment: quickstart enables it by default, the combined setup always has it on, and older multi-container deployments need `EmbeddedIdP.Enabled` set to `true` in `management.json`.
|
||||
|
||||
**Confirm**: The Identity Providers tab appears under Settings.
|
||||
|
||||
## Users not syncing from SSO provider
|
||||
|
||||
**Symptom**: Users who authenticate via SSO don't appear in the user list.
|
||||
|
||||
**Likely causes and fixes** (most common first):
|
||||
|
||||
1. **They haven't logged in yet.** Users appear after their first successful login, not when the connector is configured.
|
||||
2. **The SSO flow isn't completing.** Confirm the user lands on the dashboard after authenticating.
|
||||
3. **Token validation is failing.** Review the Management logs for token errors.
|
||||
|
||||
**Confirm**: After a successful SSO login, the user appears in the dashboard user list.
|
||||
@@ -0,0 +1,71 @@
|
||||
export const description =
|
||||
"Self-hosted NetBird installation script issues: the readiness wait, DNS, certificates, Traefik, and re-running cleanly."
|
||||
|
||||
# Installation script issues
|
||||
|
||||
Problems running the self-hosted `getting-started.sh` script. For other areas, start from [Troubleshooting](/selfhosted/troubleshooting).
|
||||
|
||||
## Script hangs on "Waiting for NetBird server to become ready"
|
||||
|
||||
**Symptom**: The `getting-started.sh` script stays on the wait line for several minutes, even though netbird-server looks healthy in `docker ps`.
|
||||
|
||||
The wait check probes the OIDC endpoint through your reverse proxy, so a healthy server alone isn't enough. Something on the path from the public internet to netbird-server is broken. Check these likely causes in order, most common first:
|
||||
|
||||
**1. DNS doesn't point at the host**
|
||||
|
||||
```bash
|
||||
dig +short netbird.example.com
|
||||
```
|
||||
|
||||
The result should be the public IP of the VM running NetBird. If it's empty or wrong, the cert challenge can't complete and the probe never succeeds.
|
||||
|
||||
**2. The certificate didn't issue**
|
||||
|
||||
For options 0 (built-in Traefik) and 1 (existing Traefik), check that a real cert was issued:
|
||||
|
||||
```bash
|
||||
curl -vI https://netbird.example.com 2>&1 | grep -E "subject|issuer"
|
||||
```
|
||||
|
||||
A self-signed or default cert means ACME hasn't completed. Check Traefik logs for ACME errors. Common causes: port 80 not reachable from the internet, DNS not propagated yet, or a Let's Encrypt rate limit.
|
||||
|
||||
**3. Traefik can't read the Docker socket (option 1 only)**
|
||||
|
||||
Traefik discovers NetBird containers via Docker labels. If the socket is mounted but unreadable (old SDK, wrong path, permission denied), routes never get created. Check Traefik logs for `client version is too old` or `permission denied while trying to connect to the Docker daemon socket`.
|
||||
|
||||
**4. Traefik and NetBird aren't on the same network (option 1 only)**
|
||||
|
||||
```bash
|
||||
docker network inspect <your-network-name> | grep -A 2 Containers
|
||||
```
|
||||
|
||||
You should see both Traefik and the NetBird containers listed.
|
||||
|
||||
**5. Routing fails end to end**
|
||||
|
||||
If the four above look fine, test the full path manually:
|
||||
|
||||
```bash
|
||||
curl -k https://netbird.example.com/oauth2/.well-known/openid-configuration
|
||||
```
|
||||
|
||||
A valid response is JSON containing `"issuer"`. Anything else points to where to dig: 502 means Traefik can't reach netbird-server, 404 means the labels aren't matched, connection refused means you're not hitting Traefik at all.
|
||||
|
||||
**Confirm**: The probe succeeds and the script continues on its own. You can leave it waiting while you debug. To start over instead, stop it with Ctrl+C, run `docker compose down -v`, fix the issue, and re-run.
|
||||
|
||||
## Script fails on existing installation check
|
||||
|
||||
**Symptom**: The script exits immediately with a message about generated files already existing.
|
||||
|
||||
**Cause**: This is intentional, and protects a working setup from being overwritten.
|
||||
|
||||
**Fix**: To start fresh (this removes all NetBird data, including users and peers):
|
||||
|
||||
```bash
|
||||
docker compose down --volumes
|
||||
rm -f docker-compose.yml dashboard.env config.yaml proxy.env \
|
||||
traefik-dynamic.yaml nginx-netbird.conf caddyfile-netbird.txt \
|
||||
npm-advanced-config.txt
|
||||
```
|
||||
|
||||
**Confirm**: Re-run the script; it proceeds past the existing-installation check and starts provisioning.
|
||||
Reference in New Issue
Block a user