The rmtree error handler in BaseNode.delete() was chmod'ing the failed
path to S_IWRITE (0o200). On POSIX this strips the search permission
from directories, turning a transient deletion failure into a directory
that can no longer be traversed or deleted. The node deletion itself
silently "succeeds" (rmtree gives up once its error handler returns)
and a later project deletion then fails with EACCES. The handler also
never retried the failed operation, so it did not help on Windows
either (the platform it was written for).
The transient failure exists in practice: a concurrent MD5 checksum
computation caching its result in the node directory (e.g. a properties
request racing the deletion) can recreate a file after rmtree has
listed the directory, making the final rmdir fail with ENOTEMPTY.
- add the missing user permissions instead of replacing the whole mode,
and retry the failed unlink/rmdir
- retry the whole deletion a few times to absorb files recreated while
the directory is being deleted
- raise a ComputeError when the directory cannot be fully deleted
instead of failing silently
Only relying on the image name lets a moved tag (e.g. a newer :latest)
silently serve stale content from a compute that already has an image
under the same name. When creating a Docker node, the controller now
pins the image id (Id from the Docker daemon on the controller host)
into the create payload. A compute holding a different image under the
same tag reports the image as missing, which routes it through the
image sync added by the previous commit and re-aligns the tag.
No new template fields or database changes: the controller host daemon
remains the source of truth and the pin is resolved per creation. When
the image is not available on the controller host the pin is omitted
and behavior is unchanged (the compute pulls from the repository).
When a Docker node is created on a remote compute whose Docker daemon
does not have the image, the compute now raises ImageMissingError
instead of blindly pulling from the Docker repository. The controller
exports the image from the Docker daemon on its host (docker save
stream) and streams it to the compute which loads it, so locally built
or docker-loaded images work across computes. When the image is not
available on the controller host either, the compute is asked to pull
it from the Docker repository as a fallback.
- add a POST /docker/images/load compute endpoint that streams a
docker save tar into the Docker daemon
- let Docker.http_query pass raw (non-dict) request bodies through so
the tar can be streamed to the daemon
- drop the inline pull from DockerVM.create() and the now unused
DockerVM.pull_image wrapper
The one-shot reclaim container inherited the image's baked-in USER:
ghcr.io/nokia/srlinux runs as "user:user", so the "privileged" helper
was exactly as unprivileged as the server itself — chmod/chown on files
written by other uids (srlinux writes as a large internal uid) failed
with EPERM and node/project deletion still broke, just with a different
error. Pass --user 0:0 explicitly so the helper is root no matter what
the image declares, and fix the manual reclaim hint the same way.
Validated live on two stuck srlinux node directories (257/258
foreign-owned entries reclaimed to 0 in ~0.5 s each).
The stop-time permission pass necessarily runs before the container's
processes exit, so files written during the shutdown window (syslog
archives, trace flushes) and after any SIGKILL path stay owned by root
on the host. An unprivileged server can neither chown nor delete them,
which broke node deletion and project deletion.
Reclaim them through the only privilege door a non-root server has:
a one-shot throwaway container of the node's own image, entrypoint
overridden to the GNS3 busybox (nothing of the guest boots), chowning
the node directory back to the server user. It runs at the end of
close() — project deletion rmtrees the directory right after the nodes
close, so close must leave a clean tree — and as a retry fallback in
delete(). The helper resolves the image by its create-time ID with
--pull=never, so a stale or retagged image name cannot turn into a
registry pull attempt.