diff --git a/docs/README.md b/docs/README.md index bc3698989..ea1581292 100644 --- a/docs/README.md +++ b/docs/README.md @@ -39,6 +39,8 @@ docs/ │ ├── statistics-api.md # Aggregated statistics API for monitoring │ ├── vnc-websocket-console.md # Browser-based VNC console via WebSocket │ └── web-wireshark-business-process.md # Web Wireshark (Docker + xpra packet capture) +├── design/ # Design proposals & roadmaps (not yet implemented) +│ └── docker-image-type.md # Docker image types: vendor profile discriminator + registry ├── gns3-copilot/ # AI Copilot feature documentation │ ├── netmiko_devices.md # Netmiko supported devices (366 types) │ ├── template-based-configuration-roadmap.md # Future: template-based config with HITL @@ -84,6 +86,16 @@ Console for vendor NOS containers (SR Linux, XRd, …) whose CLI is a TUI off PI ### Cisco XRd Control Plane (`features/vendor-nos-xrd.md`) Cisco XRd as a GNS3 Docker router: vendor path + shm/device injection (`GNS3_SHM_SIZE`/`GNS3_DEVICES`), config-file injection (`extra_configs`), udev masking (`GNS3_MASK_UDEV`) so privileged systemd containers don't disturb the host, and the host-readiness check. +### IOL Images with iol-runner (`features/iol-runner-docker.md`) +Cisco CML containerized IOL (e.g. `iol-xe/iol-xe:17-18-02`) as GNS3 Docker routers: generic unix-socket NIO (`GNS3_UNIX_SOCKET_NIO` — adapters wired via AF_UNIX datagram socket pairs instead of TAP/netns) plus `IOLDockerVM` (`GNS3_IOL_RUNNER=1` — per-start config generation, `/tmp/run` preparation, stale-socket cleanup, console on PID 1 stdio). + +--- + +## Design & Roadmaps (`design/`) + +### Docker Image Types (`design/docker-image-type.md`) +Proposed `image_type` discriminator on Docker templates plus a compute-side profile registry: vendor parameters graduate from environment markers into schema-gated fields, generic-feature applicability becomes declared capability data, and the existing markers remain as a compatibility fallback. Not a new node type — vendor images are content, not mechanism. + --- ## GNS3 AI Copilot (`gns3-copilot/`) diff --git a/docs/design/docker-image-type.md b/docs/design/docker-image-type.md new file mode 100644 index 000000000..db00cc345 --- /dev/null +++ b/docs/design/docker-image-type.md @@ -0,0 +1,144 @@ +# Docker Image Types (Vendor Profiles) — Roadmap + +Status: **proposal, not implemented**. This is the agreed direction for the +next iteration of vendor Docker support. Four vendor profiles already +exist across branches and prototypes — iol-runner (this branch), XRd and +the SR Linux prototype (vendor/skip-init family), and the FRR registry +appliance — enough evidence to justify the registry. Implement it as a +follow-up to the current PR series, shaped by all four rather than by +iol-runner alone. + +## Problem + +All Docker nodes share one template type (`template_type: "docker"`) and +one template schema. Vendor images (iol-runner, SR Linux, XRd, …) are +selected and configured today through **environment markers** — free-form +strings parsed by the compute (`GNS3_IOL_RUNNER`, `GNS3_IOL_MEMORY`, +`GNS3_IOL_STARTUP_CONFIG`, `GNS3_SKIP_INIT`, `GNS3_UNIX_SOCKET_NIO`, …). +Three pains have already surfaced: + +1. **Invisible and unvalidated.** The markers do not exist in the API + schema: no OpenAPI documentation, no validation, typo or bad value + silently falls back to a default (`GNS3_IOL_MEMORY=notanumber` → 2048). +2. **The shared-schema dilemma.** Adding a vendor-specific field to the + Docker template schema pollutes every Docker template with a field only + one vendor reads; not adding it pushes everything into environment + strings. The `startup_config` discussion (ended in the + `GNS3_IOL_STARTUP_CONFIG` knob) is the canonical example — the tension + exists because there is no legitimate discriminator dimension. +3. **Generic-feature applicability is implicit.** Which generic Docker + features apply to which image is only documented in code comments: + `mac_address` is meaningless under unix-socket NIO, `/etc/network` + seeding is dead weight for skip-init images, `extra_configs` targets + under persisted volumes get a (correct-but-noisy) shadow warning. + +## Non-goal: new node types + +A `template_type`/`node_type` per vendor (`"iol"`, `"srlinux"`, …) is +explicitly rejected. `node_type` is a top-level concept: it drives compute +module routing (`/projects/{id}/{node_type}/nodes`), capability +reporting, GUI node types and link handling. N vendor types that are all +Docker underneath would multiply routing and schema surface for zero +mechanism — vendor images are *content*, not mechanism, and GNS3's +appliance/template system is where content belongs. + +## Design + +### One discriminator field on Docker templates + +```json +{ + "template_type": "docker", + "image_type": "iol-runner", + "image": "iol-xe/iol-xe:17-18-02" +} +``` + +`image_type` is optional; absent (or `"generic"`) keeps today's plain +`DockerVM` behavior and marker sniffing. The profiles that exist today, +mapped to their current mechanisms and archetypes: + +| Profile | Current mechanism | Archetype | +|---|---|---| +| FRR appliance | plain `DockerVM`: init.sh `/etc/network` + `start_command` (frrinit.sh), console on PID 1 | generic | +| XRd | `VendorDockerVM` skip-init + `GNS3_SHM_SIZE`/`GNS3_DEVICES`, `extra_configs`, udev masking | vendor skip-init | +| SR Linux (prototype) | `VendorDockerVM` skip-init + `docker_exec` console + `GNS3_INTERFACE_NAMES` | vendor skip-init | +| iol-runner | `IOLDockerVM`: unix-socket NIO, per-start config generation, NVRAM startup-config | iol-runner | + +Possible first values: `generic`, `iol-runner`, and a value for the +skip-init vendor NOS family once its common shape settles (XRd and SR +Linux may end up sharing it or splitting — that is exactly what the +registry should decide with all four in front of us). + +### A compute-side registry + +```python +IMAGE_PROFILES = { + "iol-runner": { + "class": IOLDockerVM, + "fields": ("iol_memory", "startup_config"), # schema-gated + "capabilities": {...}, # see below + }, + ... +} +``` + +* **Class selection** moves from environment sniffing + (`Docker._select_node_class`) to the field, with the existing markers + kept as a fallback for templates created before the field existed + (zero-migration compatibility). +* **Vendor parameters graduate into schema fields**, gated by pydantic + conditional validation (accepted — and validated — only when + `image_type` matches). This resolves the shared-schema dilemma properly: + the field exists, but only means something for its profile. The + environment knobs remain as the wire-level compatibility entry. +* **Capability declaration** makes generic-feature applicability data + instead of comments: + + | Capability | plain Docker | vendor skip-init | iol-runner | + |---|---|---|---| + | `mac_address` honored | yes | yes | **no** (IOL derives MACs from the app id) | + | `/etc/network` seeding | yes | no | no | + | `extra_configs` targets under volumes | shadowed | works (real-path binds) | works | + | console | telnet/http/… | `docker_exec` | PID 1 stdio (telnet) | + | startup config | `extra_configs` injection | image-specific | nvram build (`nvram_import`) | + + Consumers: error/warning quality (don't warn about inapplicable + features), the WebUI (vendor-aware template forms — a bonus, not a + driver), and documentation generation. + +The existing class hierarchy (`DockerVM` → `VendorDockerVM` → +`IOLDockerVM`) is unchanged — the registry only externalizes selection +and declaration. + +## Migration & compatibility + +* Old templates (markers only) keep working via the fallback; a one-time + optional converter can rewrite markers → `image_type` + fields. +* The controller-side materialization conventions + (`GNS3_IOL_STARTUP_CONFIG` file → `startup_config_content`, sent once) + carry over unchanged — the field version references the same content + pipeline. + +## Roadmap steps + +1. Add `image_type` to the Docker template schema (+ DB column + + Alembic migration — see the three-place rule for template fields) and + wire class selection through the registry, markers as fallback. +2. Move iol-runner knobs to gated fields (`iol_memory`, + `startup_config`), deprecating-but-supporting the env forms. +3. Introduce the capability table and use it to silence inapplicable + warnings (`mac_address`, `/etc/network`, extra-config shadowing). +4. Revisit per-vendor *schemas* (a sub-model per profile) only if a + profile grows more than a handful of fields — not before. + +## Open questions + +* Field name: `image_type` vs `vendor` vs `runner` — decide when the + second profile lands; the value vocabulary should name the *runtime + contract*, not the vendor. +* Should appliances (`.gns3a`) carry `image_type` explicitly, or should + installation keep deriving it from the appliance's environment block? +* Where the gated vendor fields live long-term: flat on the Docker + template (simple, gated) vs nested `{"image_type": ..., "settings": + {...}}` (cleaner, more schema churn). diff --git a/docs/features/iol-runner-docker.md b/docs/features/iol-runner-docker.md new file mode 100644 index 000000000..893d77493 --- /dev/null +++ b/docs/features/iol-runner-docker.md @@ -0,0 +1,218 @@ + + +> This documentation is organized by AI with reference to actual code. AI can make mistakes — please verify against the source code when in doubt. + + +# IOL Images with iol-runner (Cisco CML containerized IOL) as Docker Nodes + +## Overview + +IOL images packaged with Cisco CML's container runner — for example +`iol-xe/iol-xe:17-18-02` (IOS-XE 17.18.02 IOL in a scratch image driven by +`iol-runner`, module `virl.lab/cmd/iol-runner`) — run as first-class GNS3 +Docker router nodes with **zero changes to the image**. The integration adds +two generic server mechanisms: + +1. **Unix-socket NIO** (`GNS3_UNIX_SOCKET_NIO=1`, on `VendorDockerVM`): link + adapters through per-interface AF_UNIX datagram sockets instead of a TAP + interface moved into the container's network namespace. +2. **`IOLDockerVM`** (marker `GNS3_IOL_RUNNER=1`): generates the runner's + config file per start, prepares its runtime directory and cleans up stale + sockets — the iol-runner-specific glue on top of the vendor path. + +## How the image works + +```mermaid +graph LR + subgraph Container["scratch container (PID 1)"] + RUNNER["/iol-runner -config /config/iol-config.json -stdio"] + IOL["IOL process
(IOS-XE 17.18.02)"] + NETIOMUX["netiomux"] + SOCKETS["/tmp/s00.sock (recv)
/tmp/c00.sock (send-to path)
… one pair per interface"] + RUNNER -->|"spawn -e/-s/-m + app id"| IOL + IOL -->|"netio bus /tmp/netio<uid>/"| NETIOMUX --> SOCKETS + end + subgraph Host + UBRIDGE["uBridge bridgeN
add_nio_unix …/gns3/unixio/<node>/cNN.sock …
+ add_nio_udp (topology)"] + RTDIR["/run/user/<uid>/gns3/unixio/<node>
(bind-mounted at /tmp)"] + VOL["project-files/docker/<node>/config
(+ tmp/run, nested bind at /tmp/run)"] + end + SOCKETS <-->|"same files, two spellings
(the /tmp bind)"| RTDIR <--> UBRIDGE + UBRIDGE -.->|"iol-config.json +
persistent /tmp/run"| VOL +``` + +* **Console**: the runner muxes the IOS console onto PID 1 stdio (`-stdio` + entrypoint flag). The plain `console_type: "telnet"` attaches to it — no + `docker_exec` needed. The runner requires a TTY, which GNS3 always + allocates; without one the runner exits (`inappropriate ioctl for device`). +* **Networking**: the runner does not touch the container's network + namespace. Per interface N it creates, inside the container's `/tmp`: a + receive socket `s%02d.sock` (frames sent there are injected into guest + interface N) and a send-to path `c%02d.sock` (whoever binds it receives + the guest's frames). Frames are **raw Ethernet**, one datagram per frame — + the same two-mailbox convention uBridge's `add_nio_unix` natively speaks + (it binds the c-socket, sends to the s-socket). GNS3 bind-mounts a + per-node directory from the runtime directory + (`/run/user//gns3/unixio/`, next to the uBridge control + sockets) at the container's `/tmp`, so uBridge reaches the sockets as + plain host files: the path stays far under AF_UNIX's 107-byte `sun_path` + cap (a projects-tree node path alone exceeds it) and the directory is + owned by the server user, to whom the runner drops its privileges. This + mirrors how CML itself runs the image (`source=…/tmp,target=/tmp` in its + node definition). The directory is ephemeral and removed with the node. +* **Licensing**: the image ships a self-consistent `/etc/hostid` + `.iourc` + pair, and the runner regenerates the license from the host ID at boot — + nothing to configure. +* **Persistence**: `/tmp/run` (the IOL working directory) holds the NETMAP, + the startup-config (`config`, plain IOS format) and NVRAM (`nvram_00001`). + It is the only `/tmp` path that needs to survive: GNS3 bind-mounts the + node directory's `tmp/run/` at `/tmp/run` (nested inside the runtime-dir + bind), so the router's configuration survives stop/start and container + recreation while sockets, netio buses and runner logs stay ephemeral. + The generated config maps the runner to the server's uid/gid + (`user-id`/`group-id`), so all files it creates are owned by the server + user (no permission-fix pass needed). + +## Template + +Create the template once via `POST /v3/templates` (authenticated — see the +API docs for the auth flow), or in the Web UI under +*Edit → Preferences → Docker templates → New* with the same fields: + +```json +{ + "name": "IOS-XE 17.18.02 IOL", + "template_type": "docker", + "image": "iol-xe/iol-xe:17-18-02", + "category": "router", + "symbol": ":/symbols/router.svg", + "adapters": 2, + "console_type": "telnet", + "environment": "GNS3_IOL_RUNNER=1", + "extra_volumes": ["/config"] +} +``` + +| Field | Value | Why | +|---|---|---| +| `environment` | `GNS3_IOL_RUNNER=1` | The switch that selects `IOLDockerVM` (skip-init, unix-socket NIO, auto volumes). Optional: `GNS3_IOL_MEMORY=` (default 2048), `GNS3_IOL_STARTUP_CONFIG=` (initial configuration, see below). | +| `extra_volumes` | `["/config"]` | `/tmp/run` is auto-added. **Never add `/tmp`** — it would persist the socket directory into the projects tree and uBridge would reject the too-long AF_UNIX path. | +| `adapters` | number of 4-port units | The IOU convention: one adapter = `Ethernet0/0`–`Ethernet0/3`, two adapters add `Ethernet1/0`–`1/3`, … (8 units / 32 ports max). Ports are addressed as (adapter, port 0–3) and shown grouped in the UI. | +| `memory` | optional; `0` (default) = no cap | Unset works — Docker applies no limit. When you do set a cap, keep it at IOL memory + ~512 MB, or the cgroup OOM-killer shoots the router. | +| `console_type` | `telnet` | The runner muxes the IOS console onto PID 1 stdio; `docker_exec` is not needed. | + +### Verify + +1. The node's port list shows the grouped IOL interfaces + (`Ethernet0/0`–`Ethernet1/3` for two adapters), addressed + (adapter, port). +2. Drop a node into a project and start it — the console shows the + `Linux Unix (i686)` banner within seconds. +3. `$XDG_RUNTIME_DIR/gns3/unixio//` contains `s00.sock`… (one + pair per port). +4. A node created with a startup-config boots straight to the configured + hostname — no initial configuration dialog (interface names in the + config are `Ethernet0/0`, not `GigabitEthernet0/0`). + +## Startup configuration + +IOL Docker nodes load an initial configuration exactly like native IOU +nodes: the **template references a config file**, and every node created +from it boots with that configuration as its personal starting point. + +```json +"environment": "GNS3_IOL_RUNNER=1\nGNS3_IOL_STARTUP_CONFIG=iol-xe-base.txt" +``` + +* The file lives in the controller's configs directory (the `configs_path` + server setting, by default `~/GNS3/configs`) — the same place IOU and + VPCS base configs live. A minimal `iol-xe-base.txt` ships with the + server and is installed there on startup (never overwriting a + user-modified copy); an absolute path in the knob bypasses the directory + entirely. `%h` in the content is replaced with the node name at start. +* **Creation materializes it once**: the controller reads the file and + sends its content with the node; afterwards the knob is consumed and the + template file is never referenced again — editing it does not affect + existing nodes. The knob also disappears from the node's `environment` + (`GNS3_IOL_RUNNER=1` remains — it re-selects the node class on every + recreation), so a node whose environment shows only `GNS3_IOL_RUNNER=1` + is the normal sign that the file was found and its content delivered; + when the file is missing the knob stays in the environment (with a + server-side warning) and the lookup is retried at the next creation. +* **NVRAM is the source of truth at boot** (verified on the runner): the + content is built into the node's `tmp/run/nvram_` — IOL and IOU + share the nvram container format, so the server-side `nvram_import` + utility produces a file IOL boots from directly. Consequences: + * `write memory` survives stop/start and container recreation (a plain + restart never re-applies the startup-config over it). The first + `write memory` after a server-built NVRAM asks for a one-time + `[confirm]` (the builder stamps an IOS 15.4 version marker into the + nvram header; press return — IOS then writes its own). + * Editing a node's `startup_config_content` property (PUT) re-applies it + on the next start, overwriting what `write memory` had saved — the + explicit edit wins, as with IOU. + * Renaming a node rewrites the `hostname` line in its NVRAM, so a + duplicated/renamed node boots under its new name. +* **Resetting a node to its template config**: stop the node, delete + `project-files/docker//tmp/run/nvram_*`, then re-apply the content + (or recreate the node). + +## Server mechanisms + +| Mechanism | Where | What it does | +|---|---|---| +| `GNS3_UNIX_SOCKET_NIO=1` | `VendorDockerVM` | `_add_ubridge_connection` override: `bridge create` + `bridge add_nio_unix /c{N:02d}.sock /s{N:02d}.sock` instead of TAP + `docker move_to_ns`. No TAP allocation, no `set_mac_addr`, namespace untouched. The socket directory is bound from a per-node runtime directory unless a persisted volume already covers it. | +| `GNS3_UNIX_SOCKET_DIR=` | `VendorDockerVM` | In-container socket directory (default `/tmp`). Any image whose agent exposes the `s%02d`/`c%02d` datagram pairs can use this without the IOL specifics. | +| `GNS3_IOL_RUNNER=1` | `IOLDockerVM` (selected in the manager) | Forces skip-init + unix-socket NIO + the `/config` and `/tmp/run` volumes; on every start writes `/config/iol-config.json` (`num-eth` = adapter count, `num-serial` = 0, memory from `GNS3_IOL_MEMORY`, default 2048), creates `/tmp/run/` (the IOL process dies without it) and removes stale sockets/netio dirs from the socket directory (`tmp/run` is never touched). | +| `restart()` hardening | `IOLDockerVM` | The base `docker restart` would boot the runner on a stale config and leave uBridge wired to the previous run's sockets; reload becomes graceful stop (SIGTERM → NVRAM flush) + full start. | +| `GNS3_IOL_STARTUP_CONFIG=` | controller `Node._node_data` + `IOLDockerVM` | Startup-config plumbing (see above): the controller translates the file into `startup_config_content` on node creation (sent once); the compute builds it into `nvram_` at the next start via the IOU `nvram_import` utility and rewrites the hostname line on rename. | + +`GNS3_STOP_TIMEOUT` (default 60) controls the SIGTERM grace period on stop. +Extra iol-runner flags can be passed via `start_command`, e.g. `-keep` +(L1 keepalives) or `-debug 9` (verbose `process.log` — very useful when +diagnosing wiring issues). + +## Notes and caveats + +* **Console typing latency**: single keystrokes are echoed on a ~50 ms + cadence — the image services console input on an internal poll. + Measured directly on the container console (bypassing the GNS3 server + entirely, through the same Docker attach endpoint it uses): keystroke + echo is bimodal — 1–11 ms when the keystroke lands just before a poll + tick, ~50–100 ms when it just misses one — while a pasted line is + picked up as one batch (~1.4 ms) and command output streams + back-to-back (inter-frame gaps ~0.01 ms). A control container on the + same attach endpoint echoes in ~1 ms, so the GNS3 telnet path is not a + factor; this matches IOS's low-priority console-input polling and there + is nothing to fix server-side. Paste long commands instead of typing + them. +* **Memory sizing**: `memory` caps the whole container; the IOL process gets + `GNS3_IOL_MEMORY` (default 2048 MB). Keep container memory at IOL memory + + ~512 MB headroom or the OOM-killer will shoot the router. +* **MAC addresses**: the `mac_address` template field and per-adapter custom + MACs are ignored — IOL derives its own scheme from the node's application + ID (`aabb.cc{app}{iface}`), e.g. `aabb.cc03.0400`. The controller + allocates the ID at node creation from the upper half of the id space + (512–1022, disjoint from IOU's 1–511, limit 511 IOL Docker nodes across + opened projects sharing computes). Starting a node without an allocation + (raw compute API, pre-allocation topologies) is an error, not a fallback — + an uncoordinated ID could collide with the pool and make nodes silently + drop each other's frames as MAC loops. +* **Interface names are IOL-style `Ethernet0/0`**, not `GigabitEthernet0/0` + (4 ports per unit, matching the adapter-count granularity) — startup + configs addressing `GigabitEthernet…` are rejected by the parser. +* **Adapters**: one adapter is a 4-port unit (the IOU model): change the + count while the node is stopped; the config is regenerated on the next + start (`num-eth` = adapters × 4) and the runner creates the matching + socket set. +* **Stop before editing**: NVRAM is only flushed on a graceful stop (SIGTERM, + "cleanup done" in `process.log`); a kill loses the running-config changes + since the last `write memory`. +* **Class selection is create-time**: toggling `GNS3_IOL_RUNNER` via PUT + takes effect after a project reload (same as `docker_exec`). +* The startup-config lives at `project-files/docker//tmp/run/config`; + `extra_configs` targets under persisted volumes are warned against by the + generic create path — edit the file directly or paste via the console. diff --git a/gns3server/api/routes/compute/docker_nodes.py b/gns3server/api/routes/compute/docker_nodes.py index 63d8bea91..4c7f319c3 100644 --- a/gns3server/api/routes/compute/docker_nodes.py +++ b/gns3server/api/routes/compute/docker_nodes.py @@ -140,6 +140,7 @@ async def update_docker_node(node_data: schemas.DockerUpdate, node: DockerVM = D "extra_hosts", "extra_volumes", "extra_configs", + "startup_config_content", "memory", "cpus", ] @@ -147,7 +148,8 @@ async def update_docker_node(node_data: schemas.DockerUpdate, node: DockerVM = D changed = False node_data = jsonable_encoder(node_data, exclude_unset=True) for prop in props: - if prop in node_data and node_data[prop] != getattr(node, prop): + # hasattr: startup_config_content only exists on IOLDockerVM + if prop in node_data and hasattr(node, prop) and node_data[prop] != getattr(node, prop): setattr(node, prop, node_data[prop]) changed = True # We don't call container.update for nothing because it will restart the container @@ -284,7 +286,7 @@ async def create_docker_node_nio( """ nio = Docker.instance().create_nio(jsonable_encoder(nio_data, exclude_unset=True)) - await node.adapter_add_nio_binding(adapter_number, nio) + await node.adapter_add_nio_binding(adapter_number, nio, port_number) return nio.asdict() @@ -302,13 +304,13 @@ async def update_docker_node_nio( The port number on the Docker node is always 0. """ - nio = node.get_nio(adapter_number) + nio = node.get_nio(adapter_number, port_number) nio.filters.clear() if nio_data.filters: nio.filters = nio_data.filters nio.markers = nio_data.markers or {} nio.suspend = nio_data.suspend - await node.adapter_update_nio_binding(adapter_number, nio) + await node.adapter_update_nio_binding(adapter_number, nio, port_number) return nio.asdict() @@ -327,7 +329,7 @@ async def delete_docker_node_nio( The port number on the Docker node is always 0. """ - await node.adapter_remove_nio_binding(adapter_number) + await node.adapter_remove_nio_binding(adapter_number, port_number) @router.post( @@ -346,7 +348,7 @@ async def start_docker_node_capture( """ pcap_file_path = os.path.join(node.project.capture_working_directory(), node_capture_data.capture_file_name) - await node.start_capture(adapter_number, pcap_file_path) + await node.start_capture(adapter_number, pcap_file_path, port_number) return {"pcap_file_path": str(pcap_file_path)} @@ -365,7 +367,7 @@ async def stop_docker_node_capture( The port number on the Docker node is always 0. """ - await node.stop_capture(adapter_number) + await node.stop_capture(adapter_number, port_number) @router.get( @@ -382,7 +384,7 @@ async def stream_pcap_file( The port number on the Docker node is always 0. """ - nio = node.get_nio(adapter_number) + nio = node.get_nio(adapter_number, port_number) stream = Docker.instance().stream_pcap_file(nio, node.project.id) return StreamingResponse(stream, media_type="application/vnd.tcpdump.pcap") diff --git a/gns3server/compute/docker/__init__.py b/gns3server/compute/docker/__init__.py index ed37ca626..f4027ea59 100644 --- a/gns3server/compute/docker/__init__.py +++ b/gns3server/compute/docker/__init__.py @@ -33,6 +33,7 @@ from gns3server.utils.asyncio import locking from gns3server.compute.base_manager import BaseManager from gns3server.compute.docker.docker_vm import DockerVM from gns3server.compute.docker.vendor_docker_vm import VendorDockerVM +from gns3server.compute.docker.iol_docker_vm import IOLDockerVM from gns3server.compute.docker.docker_error import DockerError, DockerHttp304Error, DockerHttp404Error, DockerHttp409Error log = logging.getLogger(__name__) @@ -62,9 +63,21 @@ class Docker(BaseManager): self._host_checked = False def _select_node_class(self, **kwargs): - """Select the node class based on console_type.""" + """ + Select the node class based on console_type and GNS3_* environment + markers. Like console_type, the environment is fixed at node creation + time: toggling a marker via PUT takes effect after a project reload. + """ if kwargs.get("console_type") == "docker_exec": return VendorDockerVM + environment = kwargs.get("environment") or "" + for line in environment.splitlines(): + line = line.strip().rstrip(",") + if line.startswith("GNS3_IOL_RUNNER="): + return IOLDockerVM + if line.startswith("GNS3_UNIX_SOCKET_NIO="): + # Generic capability, usable without the IOL specifics. + return VendorDockerVM return DockerVM async def create_node(self, name, project_id, node_id, *args, **kwargs): diff --git a/gns3server/compute/docker/docker_vm.py b/gns3server/compute/docker/docker_vm.py index 44a89f968..b226a0f39 100644 --- a/gns3server/compute/docker/docker_vm.py +++ b/gns3server/compute/docker/docker_vm.py @@ -899,19 +899,24 @@ class DockerVM(BaseNode): await self._start_ubridge(require_privileged_access=True) for adapter_number in range(0, self.adapters): - nio = self._ethernet_adapters[adapter_number].get_nio(0) - async with self.manager.ubridge_lock: - try: - await self._add_ubridge_connection(nio, adapter_number) - except UbridgeNamespaceError: - log.error("Container %s failed to start", self.name) - await self.stop() + adapter = self._ethernet_adapters[adapter_number] + # Single-port adapters (the standard case) loop once, keeping + # the historical command sequence; multi-port adapters + # (e.g. IOL's 4-port units) get one bridge per port. + for port_number in range(0, adapter.interfaces): + nio = adapter.get_nio(port_number) + async with self.manager.ubridge_lock: + try: + await self._add_ubridge_connection(nio, adapter_number, port_number) + except UbridgeNamespaceError: + log.error("Container %s failed to start", self.name) + await self.stop() - # The container can crash soon after the start, this means we can not move the interface to the container namespace - logdata = await self._get_log() - for line in logdata.split("\n"): - log.error(line) - raise DockerError(logdata) + # The container can crash soon after the start, this means we can not move the interface to the container namespace + logdata = await self._get_log() + for line in logdata.split("\n"): + log.error(line) + raise DockerError(logdata) await self._start_console_server() @@ -1406,6 +1411,21 @@ class DockerVM(BaseNode): """ return f"eth{adapter_number}" + def _bridge_name(self, adapter_number, port_number=0): + """ + uBridge bridge name for an adapter port. Adapters with a single + port (every standard Docker node) keep the historical + "bridge{adapter}" name; multi-port adapters (e.g. IOL's 4-port + units) get one bridge per port. + + :param adapter_number: adapter number + :param port_number: port number on the adapter + """ + + if port_number: + return f"bridge{adapter_number}_{port_number}" + return f"bridge{adapter_number}" + async def _start_interface_monitor(self): """Monitor administrative state changes for this container's adapters. @@ -1558,12 +1578,14 @@ class DockerVM(BaseNode): if writer: await self._close_interface_monitor_writer(writer) - async def _add_ubridge_connection(self, nio, adapter_number): + async def _add_ubridge_connection(self, nio, adapter_number, port_number=0): """ Creates a connection in uBridge. :param nio: NIO instance or None if it's a dummy interface (if an interface is missing in ubridge you can't see it via ifconfig in the container) :param adapter_number: adapter number + :param port_number: port number on the adapter (standard Docker + adapters have a single port, so this is always 0 on the TAP path) """ try: @@ -1575,6 +1597,13 @@ class DockerVM(BaseNode): ) ) + if port_number and adapter.interfaces == 1: + raise DockerError( + "Port {port_number} doesn't exist on adapter {adapter_number} of Docker container '{name}'".format( + name=self.name, port_number=port_number, adapter_number=adapter_number + ) + ) + for index in range(4096): if f"tap-gns3-e{index}" not in psutil.net_if_addrs(): adapter.host_ifc = f"tap-gns3-e{str(index)}" @@ -1585,12 +1614,12 @@ class DockerVM(BaseNode): name=self.name, adapter_number=adapter_number ) ) - bridge_name = f"bridge{adapter_number}" + bridge_name = self._bridge_name(adapter_number, port_number) await self._ubridge_send(f"bridge create {bridge_name}") self._bridges.add(bridge_name) await self._ubridge_send( - "bridge add_nio_tap bridge{adapter_number} {hostif} off".format( - adapter_number=adapter_number, hostif=adapter.host_ifc + "bridge add_nio_tap {bridge_name} {hostif} off".format( + bridge_name=bridge_name, hostif=adapter.host_ifc ) ) @@ -1618,23 +1647,24 @@ class DockerVM(BaseNode): log.debug(f"Created adapter {adapter_number} with MAC address {mac_address} in namespace {self._namespace}") if nio: - await self._connect_nio(adapter_number, nio) - await self._set_adapter_carrier(adapter_number, not nio.suspend) + await self._connect_nio(adapter_number, nio, port_number) + await self._set_adapter_carrier(adapter_number, not nio.suspend, port_number) async def _get_namespace(self): result = await self.manager.query("GET", f"containers/{self._cid}/json") return int(result["State"]["Pid"]) - async def _set_adapter_carrier(self, adapter_number, connected): + async def _set_adapter_carrier(self, adapter_number, connected, port_number=0): """Replicate a Docker adapter's connection state on its TAP device.""" + bridge_name = self._bridge_name(adapter_number, port_number) state = "on" if connected else "off" - await self._ubridge_send(f"bridge set_nio_tap_carrier bridge{adapter_number} {state}") + await self._ubridge_send(f"bridge set_nio_tap_carrier {bridge_name} {state}") - async def _connect_nio(self, adapter_number, nio): + async def _connect_nio(self, adapter_number, nio, port_number=0): - bridge_name = f"bridge{adapter_number}" + bridge_name = self._bridge_name(adapter_number, port_number) await self._ubridge_send( "bridge add_nio_udp {bridge_name} {lport} {rhost} {rport}".format( bridge_name=bridge_name, lport=nio.lport, rhost=nio.rhost, rport=nio.rport @@ -1650,12 +1680,13 @@ class DockerVM(BaseNode): await self._ubridge_apply_filters(bridge_name, nio.filters) await self._ubridge_apply_markers(bridge_name, nio) - async def adapter_add_nio_binding(self, adapter_number, nio): + async def adapter_add_nio_binding(self, adapter_number, nio, port_number=0): """ Adds an adapter NIO binding. :param adapter_number: adapter number - :param nio: NIO instance to add to the slot/port + :param nio: NIO instance to add to the adapter/port + :param port_number: port number on the adapter (0 for single-port adapters) """ try: @@ -1667,38 +1698,47 @@ class DockerVM(BaseNode): ) ) - if self.status == "started" and self.ubridge: - await self._connect_nio(adapter_number, nio) - await self._set_adapter_carrier(adapter_number, not nio.suspend) + if not adapter.port_exists(port_number): + raise DockerError( + "Port {port_number} doesn't exist on adapter {adapter_number} of Docker container '{name}'".format( + name=self.name, port_number=port_number, adapter_number=adapter_number + ) + ) - adapter.add_nio(0, nio) + if self.status == "started" and self.ubridge: + await self._connect_nio(adapter_number, nio, port_number) + await self._set_adapter_carrier(adapter_number, not nio.suspend, port_number) + + adapter.add_nio(port_number, nio) log.debug( "Docker container '{name}' [{id}]: {nio} added to adapter {adapter_number}".format( name=self.name, id=self._id, nio=nio, adapter_number=adapter_number ) ) - async def adapter_update_nio_binding(self, adapter_number, nio): + async def adapter_update_nio_binding(self, adapter_number, nio, port_number=0): """ Update an adapter NIO binding. :param adapter_number: adapter number :param nio: NIO instance to update the adapter + :param port_number: port number on the adapter (0 for single-port adapters) """ if self.ubridge: - bridge_name = f"bridge{adapter_number}" + bridge_name = self._bridge_name(adapter_number, port_number) if bridge_name in self._bridges: await self._ubridge_apply_filters(bridge_name, nio.filters) await self._ubridge_apply_markers(bridge_name, nio) if self.status == "started": - await self._set_adapter_carrier(adapter_number, not nio.suspend) + await self._set_adapter_carrier(adapter_number, not nio.suspend, port_number) - async def adapter_remove_nio_binding(self, adapter_number): + async def adapter_remove_nio_binding(self, adapter_number, port_number=0): """ Removes an adapter NIO binding. :param adapter_number: adapter number + :param port_number: port number on the adapter (0 for single-port adapters) :returns: NIO instance """ @@ -1712,20 +1752,20 @@ class DockerVM(BaseNode): ) ) - await self.stop_capture(adapter_number) + await self.stop_capture(adapter_number, port_number) if self.ubridge: - nio = adapter.get_nio(0) - bridge_name = f"bridge{adapter_number}" + nio = adapter.get_nio(port_number) + bridge_name = self._bridge_name(adapter_number, port_number) if self.status == "started": - await self._set_adapter_carrier(adapter_number, False) + await self._set_adapter_carrier(adapter_number, False, port_number) await self._ubridge_send(f"bridge stop {bridge_name}") await self._ubridge_send( - "bridge remove_nio_udp bridge{adapter} {lport} {rhost} {rport}".format( - adapter=adapter_number, lport=nio.lport, rhost=nio.rhost, rport=nio.rport + "bridge remove_nio_udp {bridge_name} {lport} {rhost} {rport}".format( + bridge_name=bridge_name, lport=nio.lport, rhost=nio.rhost, rport=nio.rport ) ) - adapter.remove_nio(0) + adapter.remove_nio(port_number) log.debug( "Docker VM '{name}' [{id}]: {nio} removed from adapter {adapter_number}".format( @@ -1733,11 +1773,12 @@ class DockerVM(BaseNode): ) ) - def get_nio(self, adapter_number): + def get_nio(self, adapter_number, port_number=0): """ Gets an adapter NIO binding. :param adapter_number: adapter number + :param port_number: port number on the adapter (0 for single-port adapters) :returns: NIO instance """ @@ -1751,10 +1792,10 @@ class DockerVM(BaseNode): ) ) - nio = adapter.get_nio(0) + nio = adapter.get_nio(port_number) if not nio: - raise DockerError(f"Adapter {adapter_number} is not connected") + raise DockerError(f"Adapter {adapter_number} port {port_number} is not connected") return nio @@ -1800,46 +1841,49 @@ class DockerVM(BaseNode): await self.manager.pull_image(image, progress_callback=callback) - async def _start_ubridge_capture(self, adapter_number, output_file): + async def _start_ubridge_capture(self, adapter_number, output_file, port_number=0): """ Starts a packet capture in uBridge. :param adapter_number: adapter number :param output_file: PCAP destination file for the capture + :param port_number: port number on the adapter (0 for single-port adapters) """ - adapter = f"bridge{adapter_number}" + bridge_name = self._bridge_name(adapter_number, port_number) if not self.ubridge: raise DockerError("Cannot start the packet capture: uBridge is not running") - await self._ubridge_send(f'bridge start_capture {adapter} "{output_file}"') + await self._ubridge_send(f'bridge start_capture {bridge_name} "{output_file}"') - async def _stop_ubridge_capture(self, adapter_number): + async def _stop_ubridge_capture(self, adapter_number, port_number=0): """ Stops a packet capture in uBridge. :param adapter_number: adapter number + :param port_number: port number on the adapter (0 for single-port adapters) """ - adapter = f"bridge{adapter_number}" + bridge_name = self._bridge_name(adapter_number, port_number) if not self.ubridge: raise DockerError("Cannot stop the packet capture: uBridge is not running") - await self._ubridge_send(f"bridge stop_capture {adapter}") + await self._ubridge_send(f"bridge stop_capture {bridge_name}") - async def start_capture(self, adapter_number, output_file): + async def start_capture(self, adapter_number, output_file, port_number=0): """ Starts a packet capture. :param adapter_number: adapter number :param output_file: PCAP destination file for the capture + :param port_number: port number on the adapter (0 for single-port adapters) """ - nio = self.get_nio(adapter_number) + nio = self.get_nio(adapter_number, port_number) if nio.capturing: raise DockerError(f"Packet capture is already activated on adapter {adapter_number}") nio.start_packet_capture(output_file) if self.status == "started" and self.ubridge: - await self._start_ubridge_capture(adapter_number, output_file) + await self._start_ubridge_capture(adapter_number, output_file, port_number) log.debug( "Docker VM '{name}' [{id}]: starting packet capture on adapter {adapter_number}".format( @@ -1847,19 +1891,20 @@ class DockerVM(BaseNode): ) ) - async def stop_capture(self, adapter_number): + async def stop_capture(self, adapter_number, port_number=0): """ Stops a packet capture. :param adapter_number: adapter number + :param port_number: port number on the adapter (0 for single-port adapters) """ - nio = self.get_nio(adapter_number) + nio = self.get_nio(adapter_number, port_number) if not nio.capturing: return nio.stop_packet_capture() if self.status == "started" and self.ubridge: - await self._stop_ubridge_capture(adapter_number) + await self._stop_ubridge_capture(adapter_number, port_number) log.debug( "Docker VM '{name}' [{id}]: stopping packet capture on adapter {adapter_number}".format( diff --git a/gns3server/compute/docker/iol_docker_vm.py b/gns3server/compute/docker/iol_docker_vm.py new file mode 100644 index 000000000..38042de6d --- /dev/null +++ b/gns3server/compute/docker/iol_docker_vm.py @@ -0,0 +1,375 @@ +# +# Copyright (C) 2025 GNS3 Technologies Inc. +# +# This program is free software: you can redistribute it and/or modify +# it under the terms of the GNU General Public License as published by +# the Free Software Foundation, either version 3 of the License, or +# (at your option) any later version. +# +# This program is distributed in the hope that it will be useful, +# but WITHOUT ANY WARRANTY; without even the implied warranty of +# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the +# GNU General Public License for more details. +# +# You should have received a copy of the GNU General Public License +# along with this program. If not, see . + +""" +IOL (IOS on Linux) Docker container subclass. + +Supports IOL images packaged with Cisco CML's container runner +(``iol-runner``, ``virl.lab/cmd/iol-runner``), e.g. ``iol-xe/iol-xe:17-18-02``: +a scratch image whose ENTRYPOINT is ``iol-runner -config /config/iol-config.json +-stdio``. The runner generates the license, writes the NETMAP, manages NVRAM +and muxes the IOS console onto PID 1 stdio (works with the plain ``telnet`` +console type; requires a TTY, which GNS3 always allocates). + +Networking does not use the container's network namespace at all: the runner's +netiomux exposes per-interface AF_UNIX datagram sockets in the container's +``/tmp`` (``s%02d.sock`` receive, ``c%02d.sock`` send — raw Ethernet frames), +wired by the generic ``GNS3_UNIX_SOCKET_NIO`` capability of VendorDockerVM +(uBridge reaches them through a per-node runtime directory bound at /tmp — +see ``VendorDockerVM._unix_socket_host_dir``). The controller allocates the +IOL application ID (upper half of the id space, disjoint from IOU's) so that +linked nodes get distinct MACs; starting a node without an allocation is an +error, not a fallback — an uncoordinated id could collide with the pool and +blackhole traffic as a MAC loop. + +This class is selected by the ``GNS3_IOL_RUNNER=1`` environment marker. +""" + +import contextlib +import glob +import json +import logging +import os +import re +import shutil + +from gns3server.compute.adapters.ethernet_adapter import EthernetAdapter +from gns3server.compute.docker.docker_error import DockerError, DockerHttp404Error +from gns3server.compute.docker.docker_vm import DockerVM +from gns3server.compute.docker.vendor_docker_vm import VendorDockerVM +from gns3server.compute.iou.utils.iou_export import nvram_export +from gns3server.compute.iou.utils.iou_import import nvram_import + +log = logging.getLogger(__name__) + + +class IOLDockerVM(VendorDockerVM): + """ + VendorDockerVM subclass for iol-runner images. + + Extra opt-in knob (beyond the inherited vendor ones): + + * ``GNS3_IOL_MEMORY=`` — IOL router memory passed via the generated + config (default 2048). The template ``memory`` field caps the whole + container: keep it at IOL memory + ~512 MB headroom or the kernel + OOM-killer will fire. + + The marker itself forces ``GNS3_SKIP_INIT`` and the unix-socket NIO wiring, + and auto-adds the ``/config`` and ``/tmp/run`` persistent volumes, so a + template containing only ``GNS3_IOL_RUNNER=1`` is fully configured. + + Startup configuration follows the IOU model: a template may reference a + config file with the ``GNS3_IOL_STARTUP_CONFIG`` environment knob; the + controller materializes the file content into ``startup_config_content`` + when the node is created. The content is built into the node's NVRAM + (``tmp/run/nvram_``) on the next start — IOL boots straight from + NVRAM, so a plain stop/start never re-applies it and ``write memory`` + survives restarts. + """ + + _IOL_CONFIG_DIR = "/config" + _IOL_RUN_DIR = "/tmp/run" + # The runner launches IOL with a fixed 256KB nvram (-n 256) + _IOL_NVRAM_SIZE_KB = 256 + + # Payload-delivered state. Deliberately NOT initialized in + # _parse_vendor_environment(): create() re-runs that parser on every + # (re)create (a stop removes the container, so every start recreates it) + # to pick up environment changes — resetting these there would lose the + # controller-allocated application id (MACs would flip to the fallback + # hash, colliding with the allocation pool) and any pending + # startup-config delivered by a PUT. + _application_id = None + _startup_config_content = None + _startup_config_dirty = False + + def _parse_vendor_environment(self): + + super()._parse_vendor_environment() + # The image has no shell (scratch): init.sh could neither run (its + # #!/bin/sh shebang doesn't exist) nor wait for eth interfaces that + # are never created. The console is IOS itself on PID 1 stdio. + self._gns3_init = False + self._unix_socket_nio = True + self._unix_socket_dir = "/tmp" + + self._iol_memory = 2048 + if self._environment: + for _line in self._environment.splitlines(): + _line = _line.strip().rstrip(",") + if _line.startswith("GNS3_IOL_MEMORY="): + try: + memory = int(_line.split("=", 1)[1].strip()) + if memory > 0: + self._iol_memory = memory + except ValueError: + pass + + @property + def application_id(self) -> int: + """ + IOL application ID: drives interface MACs (aabb.cc{app}{iface}) and + the NVRAM file name. Allocated by the controller from the IOL Docker + half of the id space (disjoint from IOU's) — there is deliberately + no fallback: an id derived any other way could silently collide with + an allocation and blackhole traffic as a MAC loop. Starting a node + without one raises (see _prepare_iol_runtime). + """ + + return self._application_id + + @application_id.setter + def application_id(self, value) -> None: + self._application_id = int(value) + + @property + def startup_config_content(self): + """ + Startup-config content, delivered by the controller when the node is + created from a template carrying GNS3_IOL_STARTUP_CONFIG (or updated + with new content). + """ + + return self._startup_config_content + + @startup_config_content.setter + def startup_config_content(self, content): + """ + Record new startup-config content; it is built into the node's NVRAM + on the next start. IOL boots from NVRAM whenever it holds a config, so + the content is only pushed when it actually changes — a plain + stop/start never re-applies it and `write memory` survives restarts. + """ + + if not content or content == self._startup_config_content: + # An empty value is ignored: erasing the config is not supported + # (mirrors the IOU setter) and Web clients PUT "" for unset fields. + return + self._startup_config_content = content + self._startup_config_dirty = True + + def _iol_nvram_file(self) -> str: + """ + The IOL NVRAM file inside the persistent /tmp/run volume. The name + embeds the application id (nvram_00772-style, like CML). + """ + + return os.path.join(self.working_dir, "tmp", "run", f"nvram_{self.application_id:05d}") + + def _apply_pending_startup_config(self) -> None: + """ + Build the node's NVRAM from the recorded startup-config content. IOL + and IOU share the same NVRAM container format (startup-config stored + as text inside the nvram file system), so the IOU nvram_import + utility produces a file IOL boots from directly — valid config, no + initial configuration dialog. + """ + + content = self._startup_config_content.replace("%h", self._name) + nvram_file = self._iol_nvram_file() + os.makedirs(os.path.dirname(nvram_file), exist_ok=True) + try: + nvram = nvram_import(None, content.encode("utf-8"), None, self._IOL_NVRAM_SIZE_KB) + with open(nvram_file, "wb") as f: + f.write(nvram) + except (OSError, ValueError) as e: + raise DockerError(f"Could not write IOL startup-config to NVRAM of container '{self._name}': {e}") + log.debug("IOL container '%s': startup-config written to %s", self._name, nvram_file) + + @DockerVM.name.setter + def name(self, new_name): + """ + Override: keep the hostname line inside the NVRAM in sync with the + node name (IOU parity), so a renamed or duplicated node boots under + its new name. Skipped while a content change is pending — the next + start pushes the new content with the already-updated name. + """ + + if not self._startup_config_dirty and self._application_id is not None: + nvram_file = self._iol_nvram_file() + if os.path.exists(nvram_file): + try: + with open(nvram_file, "rb") as f: + startup_config, _ = nvram_export(f.read()) + if startup_config: + content = re.sub( + r"hostname .+$", + "hostname " + new_name, + startup_config.decode("utf-8", errors="replace"), + flags=re.MULTILINE, + ) + nvram = nvram_import(None, content.encode("utf-8"), None, self._IOL_NVRAM_SIZE_KB) + with open(nvram_file, "wb") as f: + f.write(nvram) + except (OSError, ValueError) as e: + log.warning(f"Could not update hostname in NVRAM of IOL container '{self._name}': {e}") + super(IOLDockerVM, IOLDockerVM).name.__set__(self, new_name) + + def asdict(self): + """ + Override: expose the recorded startup-config content (empty for nodes + created before the knob existed or reloaded from a topology). + """ + + result = super().asdict() + result["startup_config_content"] = self._startup_config_content + return result + + @DockerVM.adapters.setter + def adapters(self, adapters): + """ + Override: one IOL adapter is a 4-port unit — the IOU model. The + template's adapter count is the number of units (2 adapters = + Ethernet0/0-3 + Ethernet1/0-3); the generated config asks the runner + for adapters × 4 interfaces and links address ports as + (adapter_number, port_number 0-3). + """ + + if len(self._ethernet_adapters) == adapters: + return + + self._ethernet_adapters.clear() + for _ in range(0, adapters): + self._ethernet_adapters.append(EthernetAdapter(interfaces=4)) + + log.debug( + "IOL container '%s': number of 4-port Ethernet adapters set to %d", + self._name, adapters, + ) + + def _persistent_volume_list(self, image_info, include_network_config=True): + """ + Override: the runner requires ``/config`` (its config file, generated + below) and ``/tmp/run`` (its working directory: startup-config and + NVRAM live there — NETMAP and the netiomux sockets are ephemeral and + stay in the container's own /tmp). Auto-add both so a minimal + template cannot be misconfigured. + """ + + volumes = super()._persistent_volume_list(image_info, include_network_config) + for needed in (self._IOL_CONFIG_DIR, self._IOL_RUN_DIR): + if not any(needed == v or needed.startswith(v.rstrip("/") + "/") for v in volumes): + volumes.append(needed) + return volumes + + async def start(self): + + await self._prepare_iol_runtime() + await super().start() + + async def restart(self): + """ + Override: the base restart is a bare ``docker restart`` — the runner + would read a stale config (no adapter-count/memory refresh) and + uBridge would keep wiring to the previous run's sockets. Stop + gracefully (SIGTERM lets the runner flush NVRAM) and start again. + """ + + await self.stop(graceful=True) + await self.start() + + async def _prepare_iol_runtime(self): + """ + Regenerate the node's runtime files before the container starts: + + * ``/tmp/run/`` must exist or the IOL process dies at + boot (the runner writes NETMAP there but does not create it). + * ``/config/iol-config.json`` is rewritten on every + start so adapter-count and memory changes take effect. + * Sockets and netio bus directories left in the wiring directory by a + previous (possibly SIGKILLed) run are removed — the runner rebinds + them on boot and would fail on a stale file. + + ``tmp/run`` (startup-config, NVRAM) is never touched. Neither is + anything while the container is already running (idempotent start of + a live node: the sockets belong to the running runner). + """ + + try: + state = await self._get_container_state() + except DockerHttp404Error: + state = "stopped" + + if self._application_id is None: + raise DockerError( + f"IOL container '{self._name}' has no application ID: nodes must be " + "created through the controller (which allocates one from the pool " + "shared with IOU), or created with an explicit application_id " + "(512-1022) on the compute API. Without a coordinated ID two nodes " + "would share MACs and drop each other's frames as loops." + ) + + os.makedirs(os.path.join(self.working_dir, "tmp", "run"), exist_ok=True) + self._write_iol_config() + + if state == "running": + return + + # Pending startup-config is materialized here rather than in the + # property setter: at create-payload time the application id may not + # be final yet (the create route applies fields in schema order) and + # a running container must not have its NVRAM swapped mid-flight. + if self._startup_config_dirty: + self._apply_pending_startup_config() + self._startup_config_dirty = False + + wiring_dir = self._unix_socket_wiring_dir() + for pattern in ("s??.sock", "c??.sock"): + for stale in glob.glob(os.path.join(wiring_dir, pattern)): + with contextlib.suppress(OSError): + os.unlink(stale) + for netio_dir in glob.glob(os.path.join(wiring_dir, "netio*")): + shutil.rmtree(netio_dir, ignore_errors=True) + + def _write_iol_config(self): + """ + Write the runner's config file on the host side of the /config volume. + The runner drops to user-id/group-id after its setup, so everything it + creates is owned by the server user — which is also what lets the + (unprivileged) uBridge write into the node's socket directory. + """ + + config = { + "binary": "/binary.iol", + "memory": self._iol_memory, + "num-eth": self.adapters * 4, # every adapter is a 4-port unit + "num-serial": 0, # GNS3 docker adapters are ethernet-only + # IOL derives interface MACs from the local application ID + # (aabb.cc{app}{iface}); every node needs a distinct one or + # linked routers share MACs and drop each other's frames as + # loops. Allocated by the controller per node (upper half of + # the id space, disjoint from IOU's — CML does the same with + # its per-deployment iol_app_id). + "local-app": self.application_id, + "remote-app": 1023, # netiomux's fake peer application ID + "user-id": os.getuid(), + "group-id": os.getgid(), + } + config_file = os.path.join(self.working_dir, "config", "iol-config.json") + os.makedirs(os.path.dirname(config_file), exist_ok=True) + with open(config_file, "w", encoding="utf-8") as f: + json.dump(config, f, indent=2) + f.write("\n") + log.debug("Wrote iol-runner config for '%s': %s", self._name, config) + + async def _fix_permissions(self): + """ + Override: no-op. The generated config maps the runner to the server's + uid/gid, so no root-owned files ever appear in the volumes, and this + image has no shell for the container-side busybox pass anyway. + """ + + self._permissions_fixed = True diff --git a/gns3server/compute/docker/vendor_docker_vm.py b/gns3server/compute/docker/vendor_docker_vm.py index 6089663f2..d6157eea4 100644 --- a/gns3server/compute/docker/vendor_docker_vm.py +++ b/gns3server/compute/docker/vendor_docker_vm.py @@ -32,7 +32,9 @@ import json import logging import os import shutil +import tempfile +from gns3server.utils.asyncio import wait_for_file_creation from gns3server.utils.asyncio.telnet_server import AsyncioTelnetServer from gns3server.compute.docker.docker_vm import DockerVM from gns3server.compute.docker.docker_error import DockerError, DockerHttp304Error, DockerHttp404Error @@ -64,6 +66,23 @@ class VendorDockerVM(DockerVM): sessions on the shared exec. * ``GNS3_STOP_TIMEOUT=60`` — SIGTERM grace period in seconds when stopping the container (default 60; Docker SIGKILLs once it expires). + * ``GNS3_UNIX_SOCKET_NIO=1`` — wire adapters through AF_UNIX datagram + socket files instead of a TAP interface moved into the container's + network namespace. For images whose network agent exposes, per adapter + ``N``, a receive socket ``s%02d.sock`` and a send-to path ``c%02d.sock`` + (raw Ethernet frames, one datagram per frame) inside the container — + e.g. Cisco CML's iol-runner (see IOLDockerVM). uBridge binds + ``c{N:02d}.sock`` (its receive side) and sends to ``s{N:02d}.sock``. + Unless the directory is already covered by a persisted volume, a + per-node directory from the runtime directory is bind-mounted there — + owned by the server user (such agents drop privileges and cannot write + into a root-owned directory) and short enough for an AF_UNIX path, + which a node directory inside the projects tree exceeds. No TAP is + created, the container's network namespace is untouched and the + ``mac_address`` template field is ignored (the image's agent owns the + MAC scheme). + * ``GNS3_UNIX_SOCKET_DIR=`` — in-container directory holding the + socket files (default ``/tmp``). """ def __init__(self, *args, **kwargs): @@ -86,6 +105,8 @@ class VendorDockerVM(DockerVM): self._console_cmd = None self._console_resize = True self._stop_timeout = 60 + self._unix_socket_nio = False + self._unix_socket_dir = "/tmp" if self._environment: for _line in self._environment.splitlines(): _line = _line.strip().rstrip(",") @@ -99,6 +120,12 @@ class VendorDockerVM(DockerVM): self._console_cmd = _line.split("=", 1)[1].strip() elif _line.startswith("GNS3_CONSOLE_RESIZE="): self._console_resize = _line.split("=", 1)[1].strip().lower() not in ("0", "false", "no") + elif _line.startswith("GNS3_UNIX_SOCKET_NIO="): + self._unix_socket_nio = _line.split("=", 1)[1].strip().lower() in ("1", "true", "yes") + elif _line.startswith("GNS3_UNIX_SOCKET_DIR="): + socket_dir = _line.split("=", 1)[1].strip().rstrip("/") or "/" + if os.path.isabs(socket_dir) and ".." not in socket_dir.split("/"): + self._unix_socket_dir = socket_dir elif _line.startswith("GNS3_STOP_TIMEOUT="): try: timeout = int(_line.split("=", 1)[1].strip()) @@ -142,6 +169,22 @@ class VendorDockerVM(DockerVM): are never shadowed by an empty mount. """ binds = super()._mount_binds(image_info) + if self._unix_socket_nio: + socket_dir = self._unix_socket_dir.rstrip("/") + if not any( + v.rstrip("/") == socket_dir or socket_dir.startswith(v.rstrip("/") + "/") + for v in self._volumes + ): + # The image's network agent drops privileges before using the + # socket directory, so it must be writable by the server user: + # bind a per-node directory from the runtime directory (see + # _unix_socket_host_dir). + binds.append({ + "Type": "bind", + "Source": self._unix_socket_host_dir(), + "Target": socket_dir, + "BindOptions": {"Propagation": "rprivate"}, + }) if self._gns3_init: return binds binds = [b for b in binds if b.get("Target") != "/gns3volumes/etc/network"] @@ -284,6 +327,133 @@ class VendorDockerVM(DockerVM): return self._interface_names[adapter_number] return f"eth{adapter_number}" + def _unix_socket_host_dir(self): + """ + Host-side directory bind-mounted at the in-container socket directory + when the latter is not already covered by a persisted volume: a + per-node directory under the runtime directory, next to the uBridge + control sockets. + + AF_UNIX paths are capped at 107 bytes (sun_path minus the NUL) — a + node directory inside the projects tree (projects// + project-files/docker//tmp) alone is ~120, so uBridge would + reject the NIO with "invalid file path size". The runtime directory + path stays short no matter where the projects live, and being owned + by the server user, the image's privilege-dropping agent can write + to it. + """ + + runtime_dir = os.environ.get("XDG_RUNTIME_DIR") or tempfile.gettempdir() + host_dir = os.path.join(runtime_dir, "gns3", "unixio", self.id) + try: + os.makedirs(host_dir, mode=0o700, exist_ok=True) + except OSError as e: + raise DockerError( + f"Could not create unix-socket directory '{host_dir}' for container '{self._name}': {e}" + ) + return host_dir + + def _remove_unix_socket_host_dir(self): + """Best-effort removal of the per-node unix-socket directory.""" + + if self._unix_socket_nio: + shutil.rmtree(self._unix_socket_host_dir(), ignore_errors=True) + + def _unix_socket_wiring_dir(self): + """ + Directory referenced in the uBridge unix-NIO commands: the persisted + volume holding the socket directory, or the per-node runtime + directory bound there by _mount_binds. + """ + + socket_dir = self._unix_socket_dir.rstrip("/") + for volume in self._volumes: + if socket_dir == volume.rstrip("/") or socket_dir.startswith(volume.rstrip("/") + "/"): + return os.path.join(self.working_dir, os.path.relpath(socket_dir, "/")) + return self._unix_socket_host_dir() + + async def delete(self): + # The per-node socket directory is ephemeral; the node is not. + await super().delete() + self._remove_unix_socket_host_dir() + + async def _add_ubridge_connection(self, nio, adapter_number, port_number=0): + """ + Override: with GNS3_UNIX_SOCKET_NIO, bridge the adapter port through + the image's AF_UNIX datagram socket pair (raw Ethernet frames) + instead of a TAP interface moved into the container's network + namespace. + + Ports are addressed flat across adapters: interface index = adapter + number × ports-per-adapter + port number (single-port adapters + reduce to the adapter number). Per interface N the image's network + agent is expected to create, inside GNS3_UNIX_SOCKET_DIR: + + * ``s{N:02d}.sock`` — its receive socket; frames sent there are + injected into guest interface N; + * ``c{N:02d}.sock`` — the path it sends guest-egress frames to. + + uBridge binds the c-socket as its receive side and sends to the + s-socket (both on the host side of the socket directory — see + _unix_socket_wiring_dir). No TAP is allocated, the namespace is + untouched and guest MAC addresses are whatever the image's agent + uses. + """ + + if not self._unix_socket_nio: + return await super()._add_ubridge_connection(nio, adapter_number, port_number) + + try: + adapter = self._ethernet_adapters[adapter_number] + except IndexError: + raise DockerError( + "Adapter {adapter_number} doesn't exist on Docker container '{name}'".format( + name=self.name, adapter_number=adapter_number + ) + ) + + interface_number = adapter_number * adapter.interfaces + port_number + bridge_name = self._bridge_name(adapter_number, port_number) + try: + await self._ubridge_send(f"bridge create {bridge_name}") + self._bridges.add(bridge_name) + + wiring_dir = self._unix_socket_wiring_dir() + local_sock = os.path.join(wiring_dir, f"c{interface_number:02d}.sock") + remote_sock = os.path.join(wiring_dir, f"s{interface_number:02d}.sock") + + # A c-socket left over from a previous ubridge run would fail its bind. + with contextlib.suppress(OSError): + os.unlink(local_sock) + + # The socket appears when the container's agent finishes its interface + # setup; wait instead of silently blackholing the adapter. + try: + await wait_for_file_creation(remote_sock, timeout=30) + except asyncio.TimeoutError: + raise DockerError( + f"Socket '{remote_sock}' for adapter {adapter_number} of container " + f"'{self._name}' did not appear within 30 seconds. Check that the " + f"container's port count covers adapter {adapter_number} and that " + f"its network agent creates the per-adapter socket pair in " + f"'{self._unix_socket_dir}'." + ) + + await self._ubridge_send(f'bridge add_nio_unix {bridge_name} "{local_sock}" "{remote_sock}"') + except Exception: + # A half-wired bridge would make the next start fail with + # "bridge already exist": uBridge stops with the failed start. + await self._stop_ubridge() + raise + adapter.host_ifc = local_sock # bookkeeping / removal logging only + log.debug( + "Adapter %d port %d of container '%s' wired via unix sockets %s <-> %s", + adapter_number, port_number, self._name, local_sock, remote_sock, + ) + + if nio: + await self._connect_nio(adapter_number, nio, port_number) + def _cleanup_console_resources(self): """ Override: close the docker-exec pty socket, if any, so the next diff --git a/gns3server/configs/iol-xe-base.txt b/gns3server/configs/iol-xe-base.txt new file mode 100644 index 000000000..fb275da85 --- /dev/null +++ b/gns3server/configs/iol-xe-base.txt @@ -0,0 +1,16 @@ +! +service timestamps debug datetime msec +service timestamps log datetime msec +! +hostname %h +! +no ip domain lookup +! +ip tcp synwait-time 5 +! +line con 0 + exec-timeout 0 0 + privilege level 15 + logging synchronous +! +end diff --git a/gns3server/controller/node.py b/gns3server/controller/node.py index 8ea491b3c..c9189f8d6 100644 --- a/gns3server/controller/node.py +++ b/gns3server/controller/node.py @@ -30,6 +30,7 @@ from .controller_error import ( from .node_types import BUILTIN_NODE_TYPES from .ports.port_factory import PortFactory, StandardPortFactory, DynamipsPortFactory from ..utils.images import images_directories +from ..utils.application_id import is_iol_runner_environment from ..utils import macaddress_to_int, int_to_macaddress from ..config import Config @@ -39,6 +40,30 @@ import logging log = logging.getLogger(__name__) +def _extract_iol_startup_config_knob(environment): + """ + Split the GNS3_IOL_STARTUP_CONFIG knob out of a Docker environment + (iol-runner images): the value is a config file name, resolved by the + controller in its configs directory. + + :returns: (filename, environment without the knob line); filename is None + when the knob is absent or empty. + """ + + if not environment or "GNS3_IOL_STARTUP_CONFIG=" not in environment: + return None, environment + filename = None + kept = [] + for line in environment.splitlines(): + stripped = line.strip().rstrip(",") + if stripped.startswith("GNS3_IOL_STARTUP_CONFIG="): + if filename is None: + filename = stripped.split("=", 1)[1].strip() + else: + kept.append(line) + return filename or None, "\n".join(kept) + + class Node: # These properties are used only on controller and are not forwarded to the compute CONTROLLER_ONLY_PROPERTIES = [ @@ -572,6 +597,25 @@ class Node: data[v] = self._base_config_file_content(self._properties[k]) del data[k] del self._properties[k] # We send the file only one time + + # IOL runner (Docker) startup-config: translate the config file + # referenced by the GNS3_IOL_STARTUP_CONFIG environment knob into + # content, like the mappings above. Sent only once — afterwards + # the node's configuration lives in its NVRAM on the compute (a + # reloaded node is recreated without the knob and boots from the + # persistent NVRAM, preserving `write memory`). + if self._node_type == "docker": + filename, environment = _extract_iol_startup_config_knob(self._properties.get("environment")) + if filename: + content = self._base_config_file_content(filename) + if content: + data["startup_config_content"] = content + data["environment"] = environment + self._properties["environment"] = environment + else: + log.warning( + f"Cannot load IOL startup-config file '{filename}' for node '{self._name}': file not found" + ) data["name"] = self._name # For remote computes, convert absolute image paths to relative paths @@ -804,22 +848,35 @@ class Node: self._ports = DynamipsPortFactory(self._properties) return elif self._node_type == "docker": - for adapter_number in range(0, self._properties["adapters"]): - custom_adapter_settings = {} - if self.custom_adapters: - for custom_adapter in self.custom_adapters: - if custom_adapter["adapter_number"] == adapter_number: - custom_adapter_settings = custom_adapter - break - port_name = f"eth{adapter_number}" - port_name = custom_adapter_settings.get("port_name", port_name) - mac_address = custom_adapter_settings.get("mac_address") - if not mac_address and "mac_address" in self._properties: - mac_address = int_to_macaddress(macaddress_to_int(self._properties["mac_address"]) + adapter_number) + if is_iol_runner_environment(self._properties.get("environment")): + # IOL adapters are 4-port units (the IOU model): ports are + # Ethernet0/0-3, Ethernet1/0-3, … addressed as + # (adapter_number, port_number 0-3). + self._ports = StandardPortFactory( + self._properties, + 4, + self._first_port_name, + "Ethernet{segment0}/{port0}", + 4, + self.custom_adapters, + ) + else: + for adapter_number in range(0, self._properties["adapters"]): + custom_adapter_settings = {} + if self.custom_adapters: + for custom_adapter in self.custom_adapters: + if custom_adapter["adapter_number"] == adapter_number: + custom_adapter_settings = custom_adapter + break + port_name = f"eth{adapter_number}" + port_name = custom_adapter_settings.get("port_name", port_name) + mac_address = custom_adapter_settings.get("mac_address") + if not mac_address and "mac_address" in self._properties: + mac_address = int_to_macaddress(macaddress_to_int(self._properties["mac_address"]) + adapter_number) - port = PortFactory(port_name, 0, adapter_number, 0, "ethernet", short_name=port_name) - port.mac_address = mac_address - self._ports.append(port) + port = PortFactory(port_name, 0, adapter_number, 0, "ethernet", short_name=port_name) + port.mac_address = mac_address + self._ports.append(port) elif self._node_type in ("ethernet_switch", "ethernet_hub"): # Basic node we don't want to have adapter number port_number = 0 diff --git a/gns3server/controller/project.py b/gns3server/controller/project.py index 7f1fad683..b69a587b6 100644 --- a/gns3server/controller/project.py +++ b/gns3server/controller/project.py @@ -40,7 +40,7 @@ from .udp_link import UDPLink from .link import _UNSET from ..config import Config from ..utils.path import check_path_allowed, get_default_project_directory -from ..utils.application_id import get_next_application_id +from ..utils.application_id import get_next_application_id, is_iol_runner_environment from ..utils.asyncio.pool import Pool from ..utils.packet_filter_validation import validate_bpf_syntax from ..utils.asyncio import locking @@ -69,6 +69,21 @@ def open_required(func): return wrapper +def _is_iol_docker_kwargs(kwargs) -> bool: + """ + Whether add_node() kwargs describe an iol-runner Docker node. The + environment can arrive nested in a ``properties`` dict or as a top-level + kwarg (the template path spreads it), matching the two shapes IOU + application-id injection handles. + """ + + if "properties" in kwargs.keys(): + environment = (kwargs.get("properties") or {}).get("environment") + else: + environment = kwargs.get("environment") + return is_iol_runner_environment(environment) + + class Project: """ A project inside a controller @@ -159,7 +174,7 @@ class Project: assert self._status != "closed" self.dump() - self._iou_id_lock = asyncio.Lock() + self._application_id_lock = asyncio.Lock() # Serialise the "ensure project exists on this compute" check in # _create_node: without it, concurrent node creations all pass the # `compute not in _project_created_on_compute` check before any has @@ -657,7 +672,7 @@ class Project: self._computes.append(compute.id) if node_type == "iou": - async with self._iou_id_lock: + async with self._application_id_lock: # IOU application IDs must be allocated serially to avoid duplicates. # The lock must also cover _create_node() because get_next_application_id() # checks in-memory nodes (self._nodes), which are only registered @@ -669,6 +684,23 @@ class Project: elif "application_id" not in kwargs.keys() and not kwargs.get("properties"): kwargs["application_id"] = get_next_application_id(self._controller.projects, self._computes) node = await self._create_node(compute, name, node_id, node_type, **kwargs) + elif node_type == "docker" and _is_iol_docker_kwargs(kwargs): + # IOL Docker nodes derive interface MACs from the application ID + # exactly like IOU; they draw from the disjoint upper half of the + # id space so the two node types can neither collide nor starve + # each other. + async with self._application_id_lock: + if "properties" in kwargs.keys(): + properties = kwargs.get("properties") or {} + if "application_id" not in properties: + properties["application_id"] = get_next_application_id( + self._controller.projects, self._computes, iol_docker=True + ) + elif "application_id" not in kwargs.keys(): + kwargs["application_id"] = get_next_application_id( + self._controller.projects, self._computes, iol_docker=True + ) + node = await self._create_node(compute, name, node_id, node_type, **kwargs) else: node = await self._create_node(compute, name, node_id, node_type, **kwargs) self.emit_notification("node.created", node.asdict()) diff --git a/gns3server/schemas/compute/docker_nodes.py b/gns3server/schemas/compute/docker_nodes.py index 7b39921f6..e2023d407 100644 --- a/gns3server/schemas/compute/docker_nodes.py +++ b/gns3server/schemas/compute/docker_nodes.py @@ -26,7 +26,7 @@ class DockerBase(BaseModel): Common Docker node properties. """ - @field_validator("start_command", "environment", "extra_hosts", mode="before") + @field_validator("start_command", "environment", "extra_hosts", "startup_config_content", mode="before") @classmethod def _empty_string_to_none(cls, value): # Web clients serialize empty form fields as "" while unset values are @@ -59,6 +59,9 @@ class DockerBase(BaseModel): extra_hosts: Optional[str] = Field(None, description="Docker extra hosts (added to /etc/hosts)") extra_volumes: Optional[List[str]] = Field(None, description="Additional directories to make persistent") extra_configs: Optional[List[ExtraConfig]] = Field(None, description="Configuration files injected into the container (bind-mounted read-only)") + startup_config_content: Optional[str] = Field( + None, description="Startup-config content (IOL runner images: materialized into the node's NVRAM at start)" + ) memory: Optional[int] = Field(None, ge=0, description="Maximum amount of memory the container can use in MB") cpus: Optional[float] = Field(None, ge=0, description="Maximum amount of CPU resources the container can use") custom_adapters: Optional[List[CustomAdapter]] = Field(None, description="Custom adapters") @@ -69,7 +72,9 @@ class DockerCreate(DockerBase): Properties to create a Docker node. """ - pass + application_id: Optional[int] = Field( + None, ge=1, le=1022, description="IOL application ID for iol-runner images (allocated by the controller)" + ) class DockerUpdate(DockerBase): diff --git a/gns3server/utils/application_id.py b/gns3server/utils/application_id.py index 9c9d8b347..488fb70b8 100644 --- a/gns3server/utils/application_id.py +++ b/gns3server/utils/application_id.py @@ -13,6 +13,7 @@ # # You should have received a copy of the GNU General Public License # along with this program. If not, see . +# from gns3server.controller.controller_error import ControllerError @@ -20,13 +21,32 @@ import logging log = logging.getLogger(__name__) +# IOU draws from the lower half: uBridge's iol_bridge uses application_id + 512 +# for its netio peer endpoint. IOL Docker nodes (iol-runner images) draw from +# the upper half: their netiomux peer is the fixed id 1023. Both node types +# derive interface MACs from the id (aabb.cc{app}{iface}), so the two pools +# must stay disjoint — a shared id silently blackholes traffic between the +# nodes as a MAC loop. +IOU_APPLICATION_ID_POOL = range(1, 512) +IOL_DOCKER_APPLICATION_ID_POOL = range(512, 1023) -def get_next_application_id(projects, computes): + +def is_iol_runner_environment(environment) -> bool: + """ + Whether a Docker node environment carries the GNS3_IOL_RUNNER marker — + the same signal the compute uses to select the IOLDockerVM class. + """ + + return "GNS3_IOL_RUNNER=" in (environment or "") + + +def get_next_application_id(projects, computes, iol_docker=False): """ Calculates free application_id from given nodes :param projects: all projects managed by controller :param computes: all computes used by the project + :param iol_docker: allocate for an IOL Docker (iol-runner) node instead of IOU :raises HTTPConflict when exceeds number :return: integer first free id """ @@ -38,12 +58,26 @@ def get_next_application_id(projects, computes): if project.status == "opened": nodes.extend(list(project.nodes.values())) - used = {n.properties["application_id"] for n in nodes if n.node_type == "iou" and n.compute.id in computes} - pool = set(range(1, 512)) + if iol_docker: + used = { + n.properties["application_id"] + for n in nodes + if n.node_type == "docker" + and n.compute.id in computes + and "application_id" in n.properties + and is_iol_runner_environment(n.properties.get("environment")) + } + pool = set(IOL_DOCKER_APPLICATION_ID_POOL) + limit = "511 IOL Docker nodes" + else: + used = {n.properties["application_id"] for n in nodes if n.node_type == "iou" and n.compute.id in computes} + pool = set(IOU_APPLICATION_ID_POOL) + limit = "512 nodes" try: application_id = (pool - used).pop() return application_id except KeyError: raise ControllerError( - "Cannot create a new IOU node (limit of 512 nodes across all opened projects using the same computes)" + f"Cannot create a new {'IOL Docker' if iol_docker else 'IOU'} node " + f"(limit of {limit} across all opened projects using the same computes)" ) diff --git a/tests/api/routes/compute/test_docker_nodes.py b/tests/api/routes/compute/test_docker_nodes.py index efadd604b..b9fa79fd8 100644 --- a/tests/api/routes/compute/test_docker_nodes.py +++ b/tests/api/routes/compute/test_docker_nodes.py @@ -20,9 +20,11 @@ import pytest_asyncio from fastapi import FastAPI, status from httpx import AsyncClient -from tests.utils import asyncio_patch -from unittest.mock import patch +from tests.utils import asyncio_patch, AsyncioMagicMock +from unittest.mock import patch, MagicMock +from gns3server.compute.docker import Docker +from gns3server.compute.docker.iol_docker_vm import IOLDockerVM from gns3server.compute.project import Project pytestmark = pytest.mark.asyncio @@ -98,6 +100,46 @@ class TestDockerNodesRoutes: assert response.json()["extra_hosts"] == "test:127.0.0.1" + async def test_docker_create_iol_startup_config( + self, app: FastAPI, + compute_client: AsyncClient, + compute_project: Project, + base_params: dict + ) -> None: + """ + The controller materializes the GNS3_IOL_STARTUP_CONFIG knob into + startup_config_content; the create route must carry it onto the + IOLDockerVM (plain Docker nodes have no such property and skip it). + """ + + params = dict(base_params) + params["image"] = "iol-xe/iol-xe:17-18-02" + params["environment"] = "GNS3_IOL_RUNNER=1" + params["startup_config_content"] = "hostname %h\n" + create_response = { + "Id": "8bd8153ea8f5", + "Warnings": [], + "Config": { + "Entrypoint": ["/iol-runner", "-config", "/config/iol-config.json", "-stdio"], + "Cmd": [], + "Volumes": {}, + }, + } + seed_proc = MagicMock() + seed_proc.communicate = AsyncioMagicMock(return_value=(b"seedcid\n", b"")) + seed_proc.returncode = 0 + with asyncio_patch("gns3server.compute.docker.Docker.list_images", return_value=[{"image": "iol-xe/iol-xe"}]): + with asyncio_patch("gns3server.compute.docker.Docker.query", return_value=create_response): + with patch("asyncio.subprocess.create_subprocess_exec", return_value=seed_proc): + response = await compute_client.post( + app.url_path_for("compute:create_docker_node", project_id=compute_project.id), json=params + ) + assert response.status_code == status.HTTP_201_CREATED + assert response.json()["startup_config_content"] == "hostname %h\n" + vm = Docker.instance().get_node(response.json()["node_id"], project_id=str(compute_project.id)) + assert isinstance(vm, IOLDockerVM) + assert vm.startup_config_content == "hostname %h\n" + @pytest.mark.parametrize( "name, status_code", ( diff --git a/tests/compute/docker/test_docker_vm.py b/tests/compute/docker/test_docker_vm.py index 10336e2b1..71dbc2c39 100644 --- a/tests/compute/docker/test_docker_vm.py +++ b/tests/compute/docker/test_docker_vm.py @@ -1198,7 +1198,7 @@ async def test_start(vm, manager, free_console_port, tmpdir): await vm.start() mock_query.assert_called_with("POST", "containers/e90e34656842/start") - vm._add_ubridge_connection.assert_called_once_with(nio, 0) + vm._add_ubridge_connection.assert_called_once_with(nio, 0, 0) assert vm._start_ubridge.called assert vm._start_console.called assert vm._start_aux.called @@ -1253,7 +1253,7 @@ async def test_start_namespace_failed(vm, manager, free_console_port): await vm.start() mock_query.assert_any_call("POST", "containers/e90e34656842/start") - mock_add_ubridge_connection.assert_called_once_with(nio, 0) + mock_add_ubridge_connection.assert_called_once_with(nio, 0, 0) assert mock_start_ubridge.called assert vm.status == "stopped" diff --git a/tests/compute/docker/test_iol_docker_vm.py b/tests/compute/docker/test_iol_docker_vm.py new file mode 100644 index 000000000..2bfc4f4ec --- /dev/null +++ b/tests/compute/docker/test_iol_docker_vm.py @@ -0,0 +1,760 @@ +# +# Copyright (C) 2025 GNS3 Technologies Inc. +# +# This program is free software: you can redistribute it and/or modify +# it under the terms of the GNU General Public License as published by +# the Free Software Foundation, either version 3 of the License, or +# (at your option) any later version. +# +# This program is distributed in the hope that it will be useful, +# but WITHOUT ANY WARRANTY; without even the implied warranty of +# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the +# GNU General Public License for more details. +# +# You should have received a copy of the GNU General Public License +# along with this program. If not, see . + +""" +Tests for the IOLDockerVM subclass (Cisco CML iol-runner images, e.g. +iol-xe/iol-xe:17-18-02) and its unix-socket NIO wiring. + +Image-free: everything is asserted against generated files, parsed knobs and +the uBridge command stream. +""" + +import asyncio +import glob +import json +import os +import uuid + +import pytest +import pytest_asyncio + +from unittest.mock import patch, MagicMock, call + +from tests.utils import asyncio_patch, AsyncioMagicMock + +from gns3server.compute.docker import Docker +from gns3server.compute.docker.docker_vm import DockerVM +from gns3server.compute.docker.vendor_docker_vm import VendorDockerVM +from gns3server.compute.docker.iol_docker_vm import IOLDockerVM +from gns3server.compute.docker.docker_error import DockerError +from gns3server.compute.iou.utils.iou_export import nvram_export +from gns3server.compute.iou.utils.iou_import import nvram_import + + +# --------------------------------------------------------------------------- +# Helpers / fixtures +# --------------------------------------------------------------------------- + +IOL_ENTRYPOINT = ["/iol-runner", "-config", "/config/iol-config.json", "-stdio"] + + +def _create_response(entrypoint=None, volumes=None): + """Build the Docker /containers/create response (with image info merged).""" + return { + "Id": "e90e34656806", + "Warnings": [], + "Config": { + "Entrypoint": entrypoint or IOL_ENTRYPOINT, + "Cmd": [], + "Volumes": volumes or {}, + }, + } + + +@pytest_asyncio.fixture +async def manager(port_manager): + + m = Docker.instance() + m.port_manager = port_manager + return m + + +def _make_vm(compute_project, manager, environment="GNS3_IOL_RUNNER=1", + extra_volumes=None, adapters=4, console_type="telnet"): + """Build an IOLDockerVM with a fake cid (no create() called).""" + vm = IOLDockerVM( + "iol-xe-1", str(uuid.uuid4()), compute_project, manager, "iol-xe/iol-xe:17-18-02", + console_type=console_type, environment=environment, + extra_volumes=extra_volumes or [], adapters=adapters, + ) + vm._cid = "e90e34656842" + # mirrors the controller flow, which always delivers an allocated id + vm.application_id = 700 + return vm + + +def _mock_start(vm, state="stopped"): + """Mock everything DockerVM.start() needs besides the runtime prep.""" + vm._get_container_state = AsyncioMagicMock(return_value=state) + vm._start_ubridge = AsyncioMagicMock() + vm._get_namespace = AsyncioMagicMock(return_value=42) + vm._add_ubridge_connection = AsyncioMagicMock() + vm._start_console_server = AsyncioMagicMock() + + +def _seed_proc(stdout=b"seedcid\n", returncode=0): + proc = MagicMock() + proc.communicate = AsyncioMagicMock(return_value=(stdout, b"")) + proc.returncode = returncode + return proc + + +@pytest.fixture(autouse=True) +def runtime_dir(tmp_path, monkeypatch): + """ + Point the unix-socket runtime directory at a per-test temporary path so + the wiring/mount tests never create per-node directories in the real one. + """ + rt = tmp_path / "run" + monkeypatch.setenv("XDG_RUNTIME_DIR", str(rt)) + return rt + + +def _mock_wiring(vm): + """Mock everything _add_ubridge_connection's unix-NIO path needs.""" + vm._ubridge_hypervisor = MagicMock() + + +def _wiring_dir(vm): + """The per-node socket directory this VM wires through.""" + return os.path.join(os.environ["XDG_RUNTIME_DIR"], "gns3", "unixio", vm.id) + + +# --------------------------------------------------------------------------- +# Factory selection +# --------------------------------------------------------------------------- + +def test_factory_selects_iol_for_env_marker(manager): + + assert manager._select_node_class(console_type="telnet", + environment="GNS3_IOL_RUNNER=1") is IOLDockerVM + + +def test_factory_tolerates_whitespace_and_comma(manager): + + assert manager._select_node_class(console_type="telnet", + environment=" GNS3_IOL_RUNNER=1,\nFOO=bar") is IOLDockerVM + + +def test_factory_docker_exec_wins_over_iol_marker(manager): + + assert manager._select_node_class(console_type="docker_exec", + environment="GNS3_IOL_RUNNER=1") is VendorDockerVM + + +def test_factory_plain_environment_is_base(manager): + + assert manager._select_node_class(console_type="telnet", + environment="FOO=bar\nGNS3_BAZ=nope") is DockerVM + + +def test_factory_generic_unix_knob_selects_vendor(manager): + + assert manager._select_node_class(console_type="telnet", + environment="GNS3_UNIX_SOCKET_NIO=1") is VendorDockerVM + + +# --------------------------------------------------------------------------- +# Knob parsing +# --------------------------------------------------------------------------- + +def test_marker_forces_skip_init_and_unix_nio(compute_project, manager): + + vm = _make_vm(compute_project, manager, environment="GNS3_IOL_RUNNER=1") + assert vm._gns3_init is False + assert vm._unix_socket_nio is True + assert vm._unix_socket_dir == "/tmp" + assert vm._iol_memory == 2048 + + +def test_iol_memory_knob(compute_project, manager): + + vm = _make_vm(compute_project, manager, + environment="GNS3_IOL_RUNNER=1\nGNS3_IOL_MEMORY=4096") + assert vm._iol_memory == 4096 + + vm = _make_vm(compute_project, manager, + environment="GNS3_IOL_RUNNER=1\nGNS3_IOL_MEMORY=notanumber") + assert vm._iol_memory == 2048 + + +# --------------------------------------------------------------------------- +# create() +# --------------------------------------------------------------------------- + +@pytest.mark.asyncio +async def test_create_keeps_image_entrypoint(compute_project, manager): + + with asyncio_patch("gns3server.compute.docker.Docker.list_images", + return_value=[{"image": "iol-xe"}]): + with asyncio_patch("gns3server.compute.docker.Docker.query", + return_value=_create_response()) as mock: + with patch("asyncio.subprocess.create_subprocess_exec", + return_value=_seed_proc()): + vm = _make_vm(compute_project, manager) + await vm.create() + sent = mock.call_args.kwargs["data"] + # the iol-runner entrypoint runs as PID 1, untouched + assert sent["Entrypoint"] == IOL_ENTRYPOINT + assert sent["Cmd"] == [] + + +@pytest.mark.asyncio +async def test_create_auto_adds_config_and_tmp_run_volumes(compute_project, manager): + + with asyncio_patch("gns3server.compute.docker.Docker.list_images", + return_value=[{"image": "iol-xe"}]): + with asyncio_patch("gns3server.compute.docker.Docker.query", + return_value=_create_response()) as mock: + with patch("asyncio.subprocess.create_subprocess_exec", + return_value=_seed_proc()): + vm = _make_vm(compute_project, manager, extra_volumes=[]) + await vm.create() + sent = mock.call_args.kwargs["data"] + mounts = sent["HostConfig"]["Mounts"] + targets = [m["Target"] for m in mounts if m.get("Type") == "bind"] + # /config (runner config) and /tmp/run (startup-config + NVRAM) + # are forced and bound at their real in-container paths + # (skip-init retargeting); /tmp is the ephemeral runtime-dir + # bind holding the netiomux sockets + assert "/config" in targets + assert "/tmp/run" in targets + tmp_mounts = [m for m in mounts if m["Target"] == "/tmp"] + assert len(tmp_mounts) == 1 + assert tmp_mounts[0]["Source"] == _wiring_dir(vm) + assert not any(t.startswith("/gns3volumes/") for t in targets) + vol_env = [v for v in sent["Env"] if v.startswith("GNS3_VOLUMES=")][0] + assert "/config" in vol_env and "/tmp/run" in vol_env + + +@pytest.mark.asyncio +async def test_create_start_command_becomes_runner_flags(compute_project, manager): + + vm = _make_vm(compute_project, manager) + vm.start_command = "-keep" + with asyncio_patch("gns3server.compute.docker.Docker.list_images", + return_value=[{"image": "iol-xe"}]): + with asyncio_patch("gns3server.compute.docker.Docker.query", + return_value=_create_response()) as mock: + with patch("asyncio.subprocess.create_subprocess_exec", + return_value=_seed_proc()): + await vm.create() + sent = mock.call_args.kwargs["data"] + # start_command is the container CMD = extra iol-runner flags + assert sent["Cmd"] == ["-keep"] + + +# --------------------------------------------------------------------------- +# start() — runtime preparation +# --------------------------------------------------------------------------- + +@pytest.mark.asyncio +async def test_start_writes_iol_config(compute_project, manager): + + vm = _make_vm(compute_project, manager, adapters=4) + _mock_start(vm) + with patch("gns3server.compute.docker.Docker.install_busybox"): + with asyncio_patch("gns3server.compute.docker.Docker.query"): + await vm.start() + + with open(os.path.join(vm.working_dir, "config", "iol-config.json")) as f: + config = json.load(f) + assert config["binary"] == "/binary.iol" + assert config["num-eth"] == 16 # 4 adapters, each a 4-port unit + assert config["num-serial"] == 0 + assert config["local-app"] == 700 + assert config["remote-app"] == 1023 + assert config["memory"] == 2048 + assert config["user-id"] == os.getuid() + assert config["group-id"] == os.getgid() + assert vm.status == "started" + + +@pytest.mark.asyncio +async def test_missing_application_id_is_an_actionable_error(compute_project, manager): + # There is no fallback id on purpose: an uncoordinated one could collide + # with a pool allocation and blackhole traffic as a MAC loop. Nodes come + # through the controller, which always allocates. + vm = _make_vm(compute_project, manager) + vm._application_id = None + vm._get_container_state = AsyncioMagicMock(return_value="stopped") + + with pytest.raises(DockerError, match="no application ID"): + await vm._prepare_iol_runtime() + + +@pytest.mark.asyncio +async def test_allocated_application_id_used(compute_project, manager): + # The controller allocates the application ID (upper half of the id + # space, disjoint from IOU's); the node uses it verbatim. + vm = _make_vm(compute_project, manager) + vm.application_id = 701 + assert vm.application_id == 701 + _mock_start(vm) + with patch("gns3server.compute.docker.Docker.install_busybox"): + with asyncio_patch("gns3server.compute.docker.Docker.query"): + await vm.start() + with open(os.path.join(vm.working_dir, "config", "iol-config.json")) as f: + config = json.load(f) + assert config["local-app"] == 701 + + +@pytest.mark.asyncio +async def test_start_rewrites_config_on_adapter_change(compute_project, manager): + + vm = _make_vm(compute_project, manager, adapters=4) + _mock_start(vm) + with patch("gns3server.compute.docker.Docker.install_busybox"): + with asyncio_patch("gns3server.compute.docker.Docker.query"): + await vm.start() + + vm.adapters = 8 + _mock_start(vm) + with patch("gns3server.compute.docker.Docker.install_busybox"): + with asyncio_patch("gns3server.compute.docker.Docker.query"): + await vm.start() + + with open(os.path.join(vm.working_dir, "config", "iol-config.json")) as f: + assert json.load(f)["num-eth"] == 32 # 8 adapters × 4 ports + + +@pytest.mark.asyncio +async def test_start_creates_run_dir(compute_project, manager): + + vm = _make_vm(compute_project, manager) + _mock_start(vm) + with patch("gns3server.compute.docker.Docker.install_busybox"): + with asyncio_patch("gns3server.compute.docker.Docker.query"): + await vm.start() + + assert os.path.isdir(os.path.join(vm.working_dir, "tmp", "run")) + + +@pytest.mark.asyncio +async def test_start_cleans_stale_sockets_but_keeps_run(compute_project, manager): + + vm = _make_vm(compute_project, manager) + wiring_dir = _wiring_dir(vm) + os.makedirs(wiring_dir, exist_ok=True) + for name in ("s00.sock", "c00.sock", "s01.sock", "c01.sock"): + open(os.path.join(wiring_dir, name), "w").close() + os.makedirs(os.path.join(wiring_dir, "netio1000")) + run_dir = os.path.join(vm.working_dir, "tmp", "run") + os.makedirs(run_dir, exist_ok=True) + open(os.path.join(run_dir, "nvram_00001"), "w").close() + open(os.path.join(run_dir, "config"), "w").close() + + _mock_start(vm, state="stopped") + with patch("gns3server.compute.docker.Docker.install_busybox"): + with asyncio_patch("gns3server.compute.docker.Docker.query"): + await vm.start() + + assert glob.glob(os.path.join(wiring_dir, "s??.sock")) == [] + assert glob.glob(os.path.join(wiring_dir, "c??.sock")) == [] + assert not os.path.exists(os.path.join(wiring_dir, "netio1000")) + # the persistent runtime survives the cleanup + assert os.path.exists(os.path.join(run_dir, "nvram_00001")) + assert os.path.exists(os.path.join(run_dir, "config")) + + +@pytest.mark.asyncio +async def test_start_skips_cleanup_when_already_running(compute_project, manager): + + vm = _make_vm(compute_project, manager) + wiring_dir = _wiring_dir(vm) + os.makedirs(wiring_dir, exist_ok=True) + open(os.path.join(wiring_dir, "s00.sock"), "w").close() + + _mock_start(vm, state="running") + with patch("gns3server.compute.docker.Docker.install_busybox"): + with asyncio_patch("gns3server.compute.docker.Docker.query"): + await vm.start() + + # live runner sockets must not be deleted behind its back + assert os.path.exists(os.path.join(wiring_dir, "s00.sock")) + + +@pytest.mark.asyncio +async def test_fix_permissions_is_noop(compute_project, manager): + + vm = _make_vm(compute_project, manager) + with patch("asyncio.subprocess.create_subprocess_exec") as mock_exec: + await vm._fix_permissions() + mock_exec.assert_not_called() + assert vm._permissions_fixed is True + + +@pytest.mark.asyncio +async def test_restart_is_graceful_stop_then_start(compute_project, manager): + + vm = _make_vm(compute_project, manager) + vm.stop = AsyncioMagicMock() + vm.start = AsyncioMagicMock() + await vm.restart() + vm.stop.assert_called_once_with(graceful=True) + vm.start.assert_called_once() + + +# --------------------------------------------------------------------------- +# Wiring — unix-socket NIO +# --------------------------------------------------------------------------- + +@pytest.mark.asyncio +async def test_add_ubridge_connection_unix_wiring(compute_project, manager): + + vm = _make_vm(compute_project, manager) + _mock_wiring(vm) + wiring_dir = _wiring_dir(vm) + os.makedirs(wiring_dir, exist_ok=True) + # the runner's receive socket must exist (created by the container) + open(os.path.join(wiring_dir, "s00.sock"), "w").close() + + nio = manager.create_nio({"type": "nio_udp", "lport": 4242, "rport": 4343, "rhost": "127.0.0.1"}) + await vm._add_ubridge_connection(nio, 0) + + sent = [c for c in vm._ubridge_hypervisor.method_calls if "send" in str(c)] + flat = "\n".join(str(c) for c in sent) + local_sock = os.path.join(wiring_dir, "c00.sock") + remote_sock = os.path.join(wiring_dir, "s00.sock") + assert call.send("bridge create bridge0") in sent + assert call.send(f'bridge add_nio_unix bridge0 "{local_sock}" "{remote_sock}"') in sent + assert "add_nio_udp bridge0 4242 127.0.0.1 4343" in flat + assert "bridge start bridge0" in flat + # the TAP/namespace path must not be used at all + assert "add_nio_tap" not in flat + assert "move_to_ns" not in flat + assert "set_mac_addr" not in flat + assert vm._ethernet_adapters[0].host_ifc == local_sock + + +@pytest.mark.asyncio +async def test_add_ubridge_connection_stale_local_socket_unlinked(compute_project, manager): + + vm = _make_vm(compute_project, manager) + _mock_wiring(vm) + wiring_dir = _wiring_dir(vm) + os.makedirs(wiring_dir, exist_ok=True) + open(os.path.join(wiring_dir, "c00.sock"), "w").close() + open(os.path.join(wiring_dir, "s00.sock"), "w").close() + + await vm._add_ubridge_connection(None, 0) + assert not os.path.exists(os.path.join(wiring_dir, "c00.sock")) + + +@pytest.mark.asyncio +async def test_add_ubridge_connection_adapter_out_of_range(compute_project, manager): + + vm = _make_vm(compute_project, manager) + _mock_wiring(vm) + with pytest.raises(DockerError): + await vm._add_ubridge_connection(None, 42) + + +@pytest.mark.asyncio +async def test_add_ubridge_connection_timeout_is_actionable(compute_project, manager): + + vm = _make_vm(compute_project, manager) + _mock_wiring(vm) + vm._stop_ubridge = AsyncioMagicMock() + + async def raise_timeout(path, timeout=60): + raise asyncio.TimeoutError() + + with patch("gns3server.compute.docker.vendor_docker_vm.wait_for_file_creation", + side_effect=raise_timeout): + with pytest.raises(DockerError) as excinfo: + await vm._add_ubridge_connection(None, 0) + # the message names the adapter and the exact wiring path + assert "adapter 0" in str(excinfo.value) + assert "s00.sock" in str(excinfo.value) + # uBridge must not survive a half-wired bridge: the retry would fail + # with "bridge already exist" + assert vm._stop_ubridge.called + + +@pytest.mark.asyncio +async def test_add_ubridge_connection_without_nio_still_wires(compute_project, manager): + + vm = _make_vm(compute_project, manager) + _mock_wiring(vm) + wiring_dir = _wiring_dir(vm) + os.makedirs(wiring_dir, exist_ok=True) + open(os.path.join(wiring_dir, "s00.sock"), "w").close() + + await vm._add_ubridge_connection(None, 0) + flat = "\n".join(str(c) for c in vm._ubridge_hypervisor.method_calls) + assert "bridge create bridge0" in flat + assert "add_nio_unix" in flat + # no link yet: no UDP NIO, no bridge start (matches base semantics) + assert "add_nio_udp" not in flat + assert "bridge start" not in flat + + +# --------------------------------------------------------------------------- +# IOU-style port model — 1 adapter = 4 ethernet ports +# --------------------------------------------------------------------------- + +def test_adapters_are_four_port_units(compute_project, manager): + + vm = _make_vm(compute_project, manager, adapters=2) + assert vm.adapters == 2 + assert len(vm._ethernet_adapters) == 2 + for adapter in vm._ethernet_adapters: + assert adapter.interfaces == 4 + for port_number in range(4): + assert adapter.port_exists(port_number) + assert not adapter.port_exists(4) + + +@pytest.mark.asyncio +async def test_wiring_addresses_ports_within_adapters(compute_project, manager): + + vm = _make_vm(compute_project, manager, adapters=2) + _mock_wiring(vm) + wiring_dir = _wiring_dir(vm) + os.makedirs(wiring_dir, exist_ok=True) + # adapter 1, port 2 -> flat interface index 6 (1 * 4 + 2) + open(os.path.join(wiring_dir, "s06.sock"), "w").close() + + nio = manager.create_nio({"type": "nio_udp", "lport": 4242, "rport": 4343, "rhost": "127.0.0.1"}) + await vm._add_ubridge_connection(nio, 1, port_number=2) + + flat = "\n".join(str(c) for c in vm._ubridge_hypervisor.method_calls) + assert 'bridge add_nio_unix bridge1_2 ' in flat + assert f'"{os.path.join(wiring_dir, "c06.sock")}"' in flat + assert f'"{os.path.join(wiring_dir, "s06.sock")}"' in flat + assert "add_nio_udp bridge1_2 4242 127.0.0.1 4343" in flat + assert vm._ethernet_adapters[1].host_ifc == os.path.join(wiring_dir, "c06.sock") + + +@pytest.mark.asyncio +async def test_nio_binding_rejects_port_out_of_range(compute_project, manager): + + vm = _make_vm(compute_project, manager, adapters=2) + nio = manager.create_nio({"type": "nio_udp", "lport": 4242, "rport": 4343, "rhost": "127.0.0.1"}) + with pytest.raises(DockerError) as excinfo: + await vm.adapter_add_nio_binding(1, nio, port_number=4) + assert "Port 4" in str(excinfo.value) + + with pytest.raises(DockerError): + await vm.adapter_add_nio_binding(9, nio, port_number=0) + + +# --------------------------------------------------------------------------- +# Generic GNS3_UNIX_SOCKET_NIO knob on plain VendorDockerVM +# --------------------------------------------------------------------------- + +def test_env_unix_socket_nio_parsing(compute_project, manager): + + vm = VendorDockerVM("vendor-1", str(uuid.uuid4()), compute_project, manager, "vendor:latest", + console_type="docker_exec", + environment="GNS3_SKIP_INIT=1\nGNS3_UNIX_SOCKET_NIO=1\nGNS3_UNIX_SOCKET_DIR=/var/run/socks") + assert vm._unix_socket_nio is True + assert vm._unix_socket_dir == "/var/run/socks" + + # invalid dirs are rejected, keeping the default + vm = VendorDockerVM("vendor-1", str(uuid.uuid4()), compute_project, manager, "vendor:latest", + console_type="docker_exec", + environment="GNS3_SKIP_INIT=1\nGNS3_UNIX_SOCKET_NIO=yes\nGNS3_UNIX_SOCKET_DIR=../../etc") + assert vm._unix_socket_nio is True + assert vm._unix_socket_dir == "/tmp" + + # off by default / explicit off + vm = VendorDockerVM("vendor-1", str(uuid.uuid4()), compute_project, manager, "vendor:latest", + console_type="docker_exec", environment="GNS3_SKIP_INIT=1") + assert vm._unix_socket_nio is False + + +def test_unix_socket_dir_bound_from_runtime_dir(compute_project, manager): + + vm = VendorDockerVM("vendor-1", str(uuid.uuid4()), compute_project, manager, "vendor:latest", + console_type="docker_exec", + environment="GNS3_SKIP_INIT=1\nGNS3_UNIX_SOCKET_NIO=1", + extra_volumes=[]) + # the socket directory is an ephemeral per-node directory from the + # runtime dir, not a volume: writable by the (unprivileged) agent and + # short enough for AF_UNIX + binds = vm._mount_binds({"Config": {"Volumes": {}}}) + socket_binds = [b for b in binds if b.get("Target") == "/tmp"] + assert len(socket_binds) == 1 + assert socket_binds[0]["Type"] == "bind" + assert socket_binds[0]["Source"] == _wiring_dir(vm) + + # a socket dir already covered by a persisted volume gets no extra bind + vm = VendorDockerVM("vendor-1", str(uuid.uuid4()), compute_project, manager, "vendor:latest", + console_type="docker_exec", + environment="GNS3_SKIP_INIT=1\nGNS3_UNIX_SOCKET_NIO=1", + extra_volumes=["/tmp"]) + binds = vm._mount_binds({"Config": {"Volumes": {}}}) + assert not any(b.get("Source") == _wiring_dir(vm) for b in binds) + assert any(b.get("Target") == "/tmp" for b in binds) + + +@pytest.mark.asyncio +async def test_generic_unix_socket_dir_honored_in_wiring(compute_project, manager): + + vm = VendorDockerVM("vendor-1", str(uuid.uuid4()), compute_project, manager, "vendor:latest", + console_type="docker_exec", + environment="GNS3_SKIP_INIT=1\nGNS3_UNIX_SOCKET_NIO=1\nGNS3_UNIX_SOCKET_DIR=/var/run/socks") + _mock_wiring(vm) + wiring_dir = _wiring_dir(vm) + os.makedirs(wiring_dir, exist_ok=True) + open(os.path.join(wiring_dir, "s00.sock"), "w").close() + + await vm._add_ubridge_connection(None, 0) + flat = "\n".join(str(c) for c in vm._ubridge_hypervisor.method_calls) + assert f'"{os.path.join(wiring_dir, "s00.sock")}"' in flat + assert "add_nio_tap" not in flat + + +# --------------------------------------------------------------------------- +# Startup configuration +# --------------------------------------------------------------------------- + + +def _nvram_path(vm): + return os.path.join(vm.working_dir, "tmp", "run", f"nvram_{vm.application_id:05d}") + + +def _read_nvram(vm): + with open(_nvram_path(vm), "rb") as f: + return f.read() + + +async def _start(vm, state="stopped"): + _mock_start(vm, state=state) + with patch("gns3server.compute.docker.Docker.install_busybox"): + with asyncio_patch("gns3server.compute.docker.Docker.query"): + await vm.start() + + +@pytest.mark.asyncio +async def test_startup_config_built_into_nvram_at_start(compute_project, manager): + + vm = _make_vm(compute_project, manager) + vm.startup_config_content = "hostname %h\ninterface Ethernet0/0\n ip address dhcp\n no shutdown" + await _start(vm) + + startup, _ = nvram_export(_read_nvram(vm)) + content = startup.decode("utf-8") + # %h is substituted with the node name, like the IOU startup-config + assert f"hostname {vm.name}" in content + assert "ip address dhcp" in content + assert vm._startup_config_dirty is False # consumed + + +@pytest.mark.asyncio +async def test_plain_restart_never_reapplies_startup_config(compute_project, manager): + # IOL boots from NVRAM whenever it holds a config: a plain stop/start must + # not rebuild it or `write memory` done in the router would be lost. + vm = _make_vm(compute_project, manager) + vm.startup_config_content = "hostname RouterOne" + await _start(vm) + + # simulate `write memory`: IOS rewrites the NVRAM behind our back + with open(_nvram_path(vm), "wb") as f: + f.write(nvram_import(None, b"hostname RouterSaved\n", None, 256)) + + await _start(vm) # plain restart, no new content delivered + + startup, _ = nvram_export(_read_nvram(vm)) + assert b"hostname RouterSaved" in startup + + +@pytest.mark.asyncio +async def test_changed_content_reapplied_at_next_start(compute_project, manager): + + vm = _make_vm(compute_project, manager) + vm.startup_config_content = "hostname RouterOne" + await _start(vm) + vm.startup_config_content = "hostname RouterTwo" + assert vm._startup_config_dirty is True + await _start(vm) + + startup, _ = nvram_export(_read_nvram(vm)) + assert b"hostname RouterTwo" in startup + + +@pytest.mark.asyncio +async def test_empty_content_never_builds_nvram(compute_project, manager): + # Web clients PUT "" for unset fields: must be ignored, not treated as a + # config change (mirrors the IOU setter's erase protection). + vm = _make_vm(compute_project, manager) + vm.startup_config_content = "" + vm.startup_config_content = None + await _start(vm) + + assert not os.path.exists(_nvram_path(vm)) + + +@pytest.mark.asyncio +async def test_reput_unchanged_content_keeps_written_nvram(compute_project, manager): + # A full PUT that carries the same content again (WebUI round-trips what + # the GET returned) must not re-apply it over `write memory`. + vm = _make_vm(compute_project, manager) + vm.startup_config_content = "hostname RouterOne" + await _start(vm) + + with open(_nvram_path(vm), "wb") as f: # simulate `write memory` + f.write(nvram_import(None, b"hostname RouterSaved\n", None, 256)) + + vm.startup_config_content = "hostname RouterOne" # unchanged re-PUT + await _start(vm) + + startup, _ = nvram_export(_read_nvram(vm)) + assert b"hostname RouterSaved" in startup + + +@pytest.mark.asyncio +async def test_running_container_keeps_pending_config(compute_project, manager): + # The NVRAM is only (re)built on a genuine start of a stopped container: + # swapping it mid-flight would be lost to the next `write memory`. + vm = _make_vm(compute_project, manager) + vm.startup_config_content = "hostname RouterOne" + await _start(vm, state="running") + + assert not os.path.exists(_nvram_path(vm)) + assert vm._startup_config_dirty is True + + +@pytest.mark.asyncio +async def test_rename_rewrites_hostname_in_nvram(compute_project, manager): + + vm = _make_vm(compute_project, manager) + vm.startup_config_content = "hostname %h" + await _start(vm) + + vm.name = "renamed-1" # IOU parity: the node boots under its new name + + startup, _ = nvram_export(_read_nvram(vm)) + assert b"hostname renamed-1" in startup + + +def test_asdict_exposes_startup_config_content(compute_project, manager): + + vm = _make_vm(compute_project, manager) + assert vm.asdict()["startup_config_content"] is None + vm.startup_config_content = "hostname RouterOne" + assert vm.asdict()["startup_config_content"] == "hostname RouterOne" + + +def test_reparse_on_recreate_keeps_payload_state(compute_project, manager): + # create() re-runs _parse_vendor_environment on every (re)create (a stop + # removes the container, so every start recreates it): payload-delivered + # state must survive the re-parse. Losing the allocated application id + # would flip MACs to the fallback hash (colliding with the allocation + # pool); losing the startup-config would drop pending PUT edits. + vm = _make_vm(compute_project, manager) + vm.application_id = 700 + vm.startup_config_content = "hostname RouterOne" + + vm._parse_vendor_environment() + + assert vm.application_id == 700 + assert vm.startup_config_content == "hostname RouterOne" + assert vm._startup_config_dirty is True + # environment-derived state still re-derives + assert vm._gns3_init is False and vm._unix_socket_nio is True diff --git a/tests/controller/test_controller.py b/tests/controller/test_controller.py index dad07581e..412a9b019 100644 --- a/tests/controller/test_controller.py +++ b/tests/controller/test_controller.py @@ -462,6 +462,9 @@ async def test_install_base_configs(controller, config, tmpdir): await controller._install_base_configs() assert os.path.exists(str(tmpdir / 'iou_l3_base_startup-config.txt')) + # the IOL docker base config ships with the server too (referenced by the + # GNS3_IOL_STARTUP_CONFIG knob in the documented template) + assert os.path.exists(str(tmpdir / 'iol-xe-base.txt')) # Check is the file has not been overwritten with open(str(tmpdir / 'iou_l2_base_startup-config.txt')) as f: diff --git a/tests/controller/test_node.py b/tests/controller/test_node.py index a20405637..e4ed15bd6 100644 --- a/tests/controller/test_node.py +++ b/tests/controller/test_node.py @@ -46,6 +46,85 @@ def node(compute, project): return node +def test_docker_iol_ports_grouped_in_four_port_units(compute, project): + """ + IOL docker nodes (GNS3_IOL_RUNNER) model adapters as 4-port units: ports + are Ethernet0/0-3, Ethernet1/0-3, … addressed (adapter, port 0-3), like + the native IOU node type. Plain docker nodes keep the flat eth naming. + """ + + node = Node(project, compute, "iol", + node_id=str(uuid.uuid4()), + node_type="docker", + properties={"adapters": 2, "environment": "GNS3_IOL_RUNNER=1"}) + ports = node.ports + assert [p.asdict()["name"] for p in ports] == [ + "Ethernet0/0", "Ethernet0/1", "Ethernet0/2", "Ethernet0/3", + "Ethernet1/0", "Ethernet1/1", "Ethernet1/2", "Ethernet1/3", + ] + assert (ports[5].adapter_number, ports[5].port_number) == (1, 1) + assert ports[5].short_name == "e1/1" + + node = Node(project, compute, "web", + node_id=str(uuid.uuid4()), + node_type="docker", + properties={"adapters": 2}) + assert [p.asdict()["name"] for p in node.ports] == ["eth0", "eth1"] + + +def test_docker_iol_startup_config_materialized_once(compute, project, monkeypatch): + """ + The GNS3_IOL_STARTUP_CONFIG environment knob references a config file in + the controller's configs directory: like the IOU startup_config mapping, + its content is sent to the compute exactly once (the knob is consumed) — + afterwards the node's configuration lives in its NVRAM on the compute. + """ + + node = Node(project, compute, "iol", + node_id=str(uuid.uuid4()), + node_type="docker", + properties={"environment": "GNS3_IOL_RUNNER=1\nGNS3_IOL_STARTUP_CONFIG=my-iol-config.txt"}) + monkeypatch.setattr(node, "_base_config_file_content", lambda path: "hostname %h\n") + + data = node._node_data() + assert data["startup_config_content"] == "hostname %h\n" + assert data["environment"] == "GNS3_IOL_RUNNER=1" + assert "GNS3_IOL_STARTUP_CONFIG" not in node.properties["environment"] + + # sent only once: a later sync (e.g. project reload) carries neither + data = node._node_data() + assert "startup_config_content" not in data + + +def test_docker_iol_startup_config_missing_file_keeps_knob(compute, project, monkeypatch): + + node = Node(project, compute, "iol", + node_id=str(uuid.uuid4()), + node_type="docker", + properties={"environment": "GNS3_IOL_RUNNER=1\nGNS3_IOL_STARTUP_CONFIG=gone.txt"}) + monkeypatch.setattr(node, "_base_config_file_content", lambda path: None) + + data = node._node_data() + assert "startup_config_content" not in data + # the knob is kept so the config still applies once the file shows up + assert "GNS3_IOL_STARTUP_CONFIG" in node.properties["environment"] + + +def test_extract_iol_startup_config_knob_variants(): + + from gns3server.controller.node import _extract_iol_startup_config_knob + + assert _extract_iol_startup_config_knob(None) == (None, None) + assert _extract_iol_startup_config_knob("GNS3_IOL_RUNNER=1") == (None, "GNS3_IOL_RUNNER=1") + # whitespace and trailing commas are tolerated, like the compute env parsing + filename, environment = _extract_iol_startup_config_knob("GNS3_IOL_RUNNER=1,\n GNS3_IOL_STARTUP_CONFIG=cfg.txt ,") + assert filename == "cfg.txt" + assert environment == "GNS3_IOL_RUNNER=1," + # an empty value is treated as absent + filename, environment = _extract_iol_startup_config_knob("GNS3_IOL_STARTUP_CONFIG=") + assert filename is None + + def test_name(compute, project): """ If node use a name template generate names diff --git a/tests/controller/test_project.py b/tests/controller/test_project.py index 31a4a2e7a..373d05e10 100644 --- a/tests/controller/test_project.py +++ b/tests/controller/test_project.py @@ -296,6 +296,47 @@ async def test_add_node_iou(controller): assert node3.properties["application_id"] == 3 +@pytest.mark.asyncio +async def test_add_node_iol_docker(controller): + """ + IOL Docker nodes (GNS3_IOL_RUNNER marker) get an application ID from the + upper half of the id space, disjoint from IOU's lower half + """ + + compute = MagicMock() + compute.id = "local" + project = await controller.add_project(project_id=str(uuid.uuid4()), name="test1") + project.emit_notification = MagicMock() + + response = MagicMock() + compute.post = AsyncioMagicMock(return_value=response) + + # template shape: environment as a top-level kwarg + node1 = await project.add_node( + compute, "iol1", None, node_type="docker", image="iol-xe/iol-xe:17-18-02", environment="GNS3_IOL_RUNNER=1" + ) + # raw API shape: environment nested in properties + node2 = await project.add_node( + compute, + "iol2", + None, + node_type="docker", + properties={"image": "iol-xe/iol-xe:17-18-02", "environment": "GNS3_IOL_RUNNER=1"}, + ) + # plain docker nodes are left alone + node3 = await project.add_node( + compute, "web", None, node_type="docker", image="nginx", environment="FOO=1", adapters=1 + ) + + assert node1.properties["application_id"] == 512 + assert node2.properties["application_id"] == 513 + assert "application_id" not in node3.properties + + # IOU keeps its own pool: a subsequent IOU node still gets the lower half + node4 = await project.add_node(compute, "iou1", None, node_type="iou") + assert node4.properties["application_id"] == 1 + + @pytest.mark.asyncio async def test_add_node_iou_with_multiple_projects(controller): """