fix: host-side permission fix + SKIP_INIT volume persistence docs

Replace the container-side _fix_permissions for vendor NOS containers with a
host-side pass that walks the node's project directories directly (they are
the Docker bind-mount sources): records mode:uid:gid into .gns3_perms and
chowns to the GNS3 user. No docker exec, no container restart — the base
implementation restarts an exited container just to chown, and after the
restart the mount --bind bridge is gone so it would fix the overlay copy
instead of the host files.

The pass runs at start (after _setup_skip_init_volumes seeds and bridges the
volumes) so the controller can read project files while the node runs, and
again at stop for files written during runtime.

Update docker-exec-console.md: VendorDockerVM architecture, hook points,
class-selection factory, volume-persistence lifecycle, and new
troubleshooting entries.
This commit is contained in:
YueGuobin 2026-08-12 22:58:46 +08:00
parent 5388fd3796
commit 2f36471a55
No known key found for this signature in database
2 changed files with 215 additions and 15 deletions

View File

@ -43,12 +43,47 @@ they stay host-side configuration.
| `GNS3_INTERFACE_NAMES=mgmt0,e1-1,e1-2,e1-3` | Rename the injected interfaces in adapter order instead of the default `eth{N}`. SR Linux expects `mgmt0` + `e1-N`; without this it does not recognise its datapath. |
| `GNS3_CONSOLE_CMD=/opt/srlinux/bin/sr_cli` | Command run by the `docker_exec` console inside the container. |
## Architecture: `VendorDockerVM` subclass
All vendor-specific logic lives in a `VendorDockerVM(DockerVM)` subclass in
`gns3server/compute/docker/vendor_docker_vm.py``docker_vm.py` itself stays
on its baseline behaviour and is never touched by this feature.
`DockerVM` exposes four small extension hooks (pure refactorings, zero
behaviour change for existing nodes):
| Hook | Baseline behaviour | `VendorDockerVM` override |
|------|--------------------|---------------------------|
| `_prepare_init_and_interface_env(params)` | prepend `/gns3/init.sh`, set `GNS3_MAX_ETHERNET=eth{N-1}` | conditional init.sh (`GNS3_SKIP_INIT`), `GNS3_MAX_ETHERNET` follows the interface rename |
| `_start_console_server()` | telnet/ssh/http console dispatch | adds the `docker_exec` branch |
| `_get_container_ifname(adapter_number)` | `eth{N}` | `GNS3_INTERFACE_NAMES` lookup, fallback `eth{N}` |
| `_cleanup_console_resources()` | no-op | closes the docker-exec pty socket before restart/stop |
### Class selection
The Docker manager picks the class per node in `Docker.create_node()`
(`gns3server/compute/docker/__init__.py`):
```python
def _select_node_class(self, **kwargs):
if kwargs.get("console_type") == "docker_exec":
return VendorDockerVM
return DockerVM
```
`console_type == "docker_exec"` is the **only** trigger — every other console
type (telnet, vnc, ssh, http, …) keeps using the unmodified `DockerVM`. All
vendor features are opt-in: without the `GNS3_*` environment variables a
`VendorDockerVM` instance behaves identically to `DockerVM` (init.sh still
runs, interfaces stay `eth{N}`, the exec command defaults to `/bin/sh`), so a
regular container can use `docker_exec` too.
## The `docker_exec` console type
Setting `console_type: "docker_exec"` makes the node's primary console port run
`_start_docker_exec_console()` instead of the attach-to-PID-1 path.
### Architecture
### Console architecture
```mermaid
graph LR
@ -67,7 +102,8 @@ behaves.
### Implementation
**File**: `gns3server/compute/docker/docker_vm.py``_start_docker_exec_console()`
**File**: `gns3server/compute/docker/vendor_docker_vm.py`
`_start_docker_exec_console()`
A small subclass `_LazyExecTelnetServer(AsyncioTelnetServer)` implements the
console. Key points:
@ -88,7 +124,13 @@ console. Key points:
CLI"* otherwise), and `Env: ["TERM=xterm"]` (the TUI library needs a
recognised terminal).
3. **Hijacked raw-HTTP start.** The exec is started with
3. **while-true wrapper.** The command is wrapped in
`sh -c "while true; do <cmd>; done"` so that when the CLI exits (user types
`quit`, or the NOS's own idle timeout logs the session out), a fresh CLI
instance starts in the same pty instead of killing the shared console
session.
4. **Hijacked raw-HTTP start.** The exec is started with
`POST exec/{eid}/start` sent as a raw HTTP upgrade over the Docker unix
socket (`asyncio.open_unix_connection`), the same approach docker-py uses.
This is required because aiohttp's websocket client (`ws_connect`) is
@ -96,11 +138,11 @@ console. Key points:
upgrade succeeds (101). With `Tty:true` the response body is a raw,
non-multiplexed bidirectional pty byte stream — no frame demux needed.
4. **NAWS → exec resize.** The telnet server runs with `naws=True`; the
5. **NAWS → exec resize.** The telnet server runs with `naws=True`; the
`window_size_changed_callback` calls `POST exec/{eid}/resize?h=&w=` so the
TUI lays out for the xterm.js window size.
5. **Binary passthrough + redraw.** `binary=True` so TUI escape sequences reach
6. **Binary passthrough + redraw.** `binary=True` so TUI escape sequences reach
xterm.js intact; `echo=False` (the pty echoes). On every client (re)connect
a `Ctrl-L` (`\x0c`) is sent to the pty so a TUI that already drew its
screen for a previous client redraws for the new one (otherwise a
@ -147,9 +189,62 @@ The exec-API approach fixes all of these: a real pty (`Tty:true`), a real size
### Persistent state
For SR Linux, persist `/etc/opt/srlinux` (config / AAA users / TLS certs) by
adding it to the node's `extra_volumes`. `/var/opt/srlinux` does **not** exist
on current SR Linux images; `/var/log/srlinux` holds logs (optional).
For SR Linux, persist `/etc/opt/srlinux` (config / AAA users / TLS certs) and
`/var/log/srlinux` (logs, optional) by adding them to the node's
`extra_volumes`. The image also declares its own `VOLUME` directories
(e.g. `/opt/srlinux/appmgr`), which GNS3 persists automatically.
## Volume persistence with `GNS3_SKIP_INIT`
This is the one place where skipping init.sh changes behaviour beyond boot:
`/gns3/init.sh` normally performs the volume-persistence bridge, and without it
**nothing writes through to the host** — the container writes to its overlay
filesystem and the data is lost on stop.
The bridge (see init.sh lines 3552) has two parts:
```
host ──Docker bind mount──▶ /gns3volumes/etc/opt/srlinux (always mounted)
│ init.sh: mount --bind
/etc/opt/srlinux (where the NOS writes)
```
`VendorDockerVM` replicates this for SKIP_INIT containers:
1. **`_setup_skip_init_volumes()`** — runs once per start, right after the
container is up (`VendorDockerVM.start()`). For each persistent volume it
`docker exec`s a busybox script that:
- seeds the host directory with the container's original files on first
start (`cp -a` + `.gns3_perms` marker), exactly like init.sh;
- `mount --bind /gns3volumes<path> <path>` to bridge persistent storage
back to the in-container path — on subsequent starts the persisted data
replaces the fresh overlay content;
- restores the permissions recorded in `.gns3_perms` at the previous stop
(best-effort).
2. **Host-side `_fix_permissions()` override**`DockerVM._fix_permissions`
is container-side (busybox via `docker exec`) and restarts an exited
container just to chown; after a restart the `mount --bind` bridge is gone,
so it would fix the overlay copy and not the host files. The override
instead walks the host-side directories under the node's project directory
directly (they *are* the Docker bind-mount sources), records
`mode:uid:gid:path` into `.gns3_perms` and chowns to the GNS3 user —
no running container required, no restart. It runs both at start (so the
controller can read project files while the node runs) and at stop.
> Rootful-Docker assumption: the `.gns3_perms` uid/gid values are recorded from
> the host's view. With rootful Docker (no userns remap) in-container and host
> ids coincide, so restore semantics are identical to init.sh's. This would
> need revisiting for userns-remapped daemons.
### Lifecycle summary
| Phase | Normal Docker node | `VendorDockerVM` + `GNS3_SKIP_INIT` |
|-------|--------------------|--------------------------------------|
| start | init.sh seeds + bind-mounts + restores perms (in-container, before the app starts) | `docker exec` after start: seed + bind-mount + restore perms; then host-side chown |
| stop | container-side `_fix_permissions` (restarts an exited container) | host-side `_fix_permissions` (no container needed) |
| volume config | identical `_mount_binds` (host → `/gns3volumes<path>`) | identical |
## Troubleshooting
@ -180,23 +275,51 @@ on current SR Linux images; `/var/log/srlinux` holds logs (optional).
network-instance before ping works. This is SR Linux behaviour, not a GNS3
issue.
**7. "Session has been idle, will logout in 300 seconds" → Connection closed**
- SR Linux's own CLI idle timeout. The while-true wrapper restarts the CLI
automatically, but to keep a permanent session disable the timeout in the
CLI: `enter candidate``/system cli idle-timeout disable``commit now`.
**8. Controller logs `Permission denied` reading files under the node's
project directory while the node runs**
- Root-written files inside a persistent volume. The host-side
`_fix_permissions` pass runs at start (fixes the seeded files) and at stop;
files created by the container *during* runtime become readable after the
next stop.
**9. Persistent volume empty on the host after `save` + stop**
- Ensure `GNS3_SKIP_INIT=1` is set (so the host-side bridge path is taken) and
the volume path is in `extra_volumes`; check the compute log for
`Volume '<path>' bound to persistent storage`.
## Limitations
1. **Shared session (broadcast).** All console clients share one CLI session
and can see each other's input — identical to GNS3's existing primary
console model. There is no per-client independent session.
2. **`reset_console` not wired.** The console-reset action only handles
`telnet`/`ssh`; it is a no-op for `docker_exec` (non-blocking; reconnect
works fine).
`telnet`/`ssh`; it is a no-op for `docker_exec` (non-blocking; reconnect
works fine).
3. **Prototype knobs.** `GNS3_SKIP_INIT` / `GNS3_INTERFACE_NAMES` /
`GNS3_CONSOLE_CMD` are environment-driven; they are not yet first-class node
schema fields and are not declared in the appliance (`gns3a`) schema.
`GNS3_CONSOLE_CMD` are environment-driven; they are not yet first-class node
schema fields and are not declared in the appliance (`gns3a`) schema.
4. **Rootful-Docker assumption** for the host-side `.gns3_perms` recording
(see the volume-persistence section).
5. **Post-boot volume bridge.** The bind-mount bridge is established after the
vendor entrypoint has started (init.sh would do it before). A NOS that
strictly requires its persisted files at its very first read may need a
different boot arrangement.
## References
- `gns3server/compute/docker/docker_vm.py``_start_docker_exec_console`,
`_LazyExecTelnetServer`, `_add_ubridge_connection` (interface rename),
`create()` (env parsing, skip-init, `GNS3_MAX_ETHERNET`).
- `gns3server/compute/docker/vendor_docker_vm.py``VendorDockerVM`:
`_start_docker_exec_console`, `_LazyExecTelnetServer`,
`_setup_skip_init_volumes`, host-side `_fix_permissions`, `start()`.
- `gns3server/compute/docker/docker_vm.py``DockerVM` extension hooks
(`_prepare_init_and_interface_env`, `_start_console_server`,
`_get_container_ifname`, `_cleanup_console_resources`).
- `gns3server/compute/docker/__init__.py``Docker._select_node_class` /
`create_node` factory.
- `gns3server/compute/base_node.py` — console WebSocket guard.
- `gns3server/schemas/common.py``ConsoleType.docker_exec`.
- containerlab `nodes/srl/srl.go` — reference for SR Linux launch command and
@ -206,4 +329,5 @@ on current SR Linux images; `/var/log/srlinux` holds logs (optional).
| Version | Date | Changes |
|---------|------|---------|
| 1.1 | 2026-08-12 | Refactor: vendor logic extracted from `DockerVM` into `VendorDockerVM` subclass with 4 hook points + class-selection factory. Add SKIP_INIT volume persistence (`_setup_skip_init_volumes` + host-side `_fix_permissions`) and lifecycle comparison. Add troubleshooting entries for idle timeout and permission-denied files. |
| 1.0 | 2026-08-12 | Initial documentation of the `docker_exec` console and vendor NOS knobs. |

View File

@ -29,6 +29,8 @@ container behaves identically to DockerVM.
import asyncio
import json
import logging
import os
import stat
from gns3server.utils.asyncio.telnet_server import AsyncioTelnetServer
from gns3server.compute.docker.docker_vm import DockerVM
@ -119,6 +121,80 @@ class VendorDockerVM(DockerVM):
await super().start()
if self.status == "started" and not self._gns3_init:
await self._setup_skip_init_volumes()
# Fix host-side ownership of the seeded volume right away so the
# controller can read project files while the node runs. Reset the
# "fixed" flag afterwards: files written by the container during
# runtime still need the stop-time pass.
await self._fix_permissions()
self._permissions_fixed = False
async def _fix_permissions(self):
"""
Host-side override of DockerVM._fix_permissions for SKIP_INIT
containers. The persistent volumes are Docker bind mounts of
directories under the node's project directory, so ownership is fixed
directly on the host no docker exec, no container restart required
(the base implementation restarts an exited container just to chown,
which is wasteful for vendor NOS images).
Two passes per volume, mirroring the base/busybox behaviour:
1. record each entry's container-visible mode/uid/gid into
`.gns3_perms` (same `mode:uid:gid:path` format init.sh consumes,
paths are in-container absolute so the restore inside the
container resolves them);
2. chmod u+rX + chown to the host user so the GNS3 process can read
and delete files from the project directory.
"""
uid, gid = os.getuid(), os.getgid()
for volume in self._volumes:
path = os.path.join(self.working_dir, os.path.relpath(volume, "/"))
if not os.path.isdir(path):
continue
def onerror(exc):
log.debug("Could not walk '%s' for container '%s': %s", exc.filename, self._name, exc)
# 1. record container-visible permissions for restore at next start
try:
with open(os.path.join(path, ".gns3_perms"), "w") as perms_file:
for root, dirs, files in os.walk(path, onerror=onerror):
for entry in dirs + files:
entry_path = os.path.join(root, entry)
try:
st = os.lstat(entry_path)
except OSError:
continue
container_path = os.path.join(volume, os.path.relpath(entry_path, path))
perms_file.write(
f"{stat.S_IMODE(st.st_mode):o}:{st.st_uid}:{st.st_gid}:{container_path}\n"
)
except OSError as e:
log.warning(
"Could not record permissions for '%s' on container '%s': %s", path, self._name, e
)
continue
# 2. chmod u+rX + chown to the host user
for root, dirs, files in os.walk(path, onerror=onerror):
for entry in dirs + files:
entry_path = os.path.join(root, entry)
try:
st = os.lstat(entry_path)
is_link = stat.S_ISLNK(st.st_mode)
if not is_link:
mode = stat.S_IMODE(st.st_mode)
new_mode = mode | 0o400 # u+r
if stat.S_ISDIR(st.st_mode) or (mode & 0o111): # u+X
new_mode |= 0o100
os.chmod(entry_path, new_mode)
os.lchown(entry_path, uid, gid)
except OSError as e:
log.debug(
"Could not fix permissions on '%s' for container '%s': %s",
entry_path, self._name, e,
)
self._permissions_fixed = True
async def _setup_skip_init_volumes(self):
"""