Guobin Yue f07d21c511
feat(web-wireshark): add implementation plan for script-driven (#2666)
* feat(web-wireshark): add implementation plan for script-driven integration

Add comprehensive implementation plan for Web Wireshark integration using script-driven approach. The plan outlines:

- Background and motivation for script-based Web Wireshark integration
- Core workflow from API request to WebSocket proxy connection
- Management script architecture with WebWiresharkManager class
- Docker container configuration and resource management
- Performance analysis and optimization recommendations
- Network setup and management procedures

Key features include:
- JWT token extraction from Authorization header
- Automatic GNS3 server URL detection
- Container lifecycle management per project
- Xpra session isolation per link
- WebSocket proxy integration through gns3server
- Resource monitoring and scaling guidelines

The implementation enables users to start Web Wireshark sessions via POST requests with `wireshark: true` parameter, providing a unified web-based packet capture interface.

* feat(web-wireshark): implement script-driven Web Wireshark integration

Add complete Web Wireshark container integration with script-driven architecture:

- Create manage_wireshark.py script for Docker container management
- Add LinkCapture schema with wireshark boolean field
- Update Link class with wireshark and jwt_token parameters
- Implement WebSocket proxy endpoint for xpra HTML5 client
- Add cleanup logic in Project class for container lifecycle

Key features:
- Script-driven container and xpra session management
- JWT token extraction from Authorization header
- GNS3 URL auto-detection with Controller/Config fallback
- WebSocket proxy for unified access through gns3server
- Proper cleanup on project close/delete

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(web-wireshark): use aiohttp instead of docker SDK

Replace docker-py SDK with direct Docker HTTP API calls using aiohttp,
following GNS3's existing architecture pattern in gns3server/compute/docker/.

Changes:
- Create DockerHTTPClient class using aiohttp + Unix socket
- Implement all Docker operations as async methods
- Remove dependency on docker Python SDK
- Use Docker API v1.44 via /var/run/docker.sock

Benefits:
- No additional dependencies required
- Consistent with GNS3's async architecture
- Better integration with existing codebase

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(web-wireshark): add dynamic Docker API version detection

Implement dynamic Docker API version detection following GNS3's pattern
in gns3server/compute/docker/__init__.py.

Changes:
- Add DOCKER_MINIMUM_API_VERSION and DOCKER_PREFERRED_API_VERSION constants
- Remove hardcoded "v1.44" prefix (now using "1.44" format)
- Add _check_connection() method to detect Docker daemon version
- Dynamically select API version based on daemon capabilities
- Add proper error handling for version mismatches

Behavior:
- Initialize with minimum API version (1.40)
- On first connection, detect Docker daemon API version
- Use preferred API version (1.44) if supported
- Fall back to daemon's minimum API version if needed
- Raise error if daemon version is too old

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(web-wireshark): add optional --verbose logging parameter

Remove hardcoded logging.basicConfig() and add --verbose parameter
to control log output, following GNS3's logging pattern.

Changes:
- Remove logging.basicConfig(level=logging.INFO) from module level
- Add --verbose/-v command line parameter
- Configure logging only when --verbose is specified
- Use structured log format with timestamp when verbose

Behavior:
- When called by GNS3: uses GNS3's logging configuration (default)
- When run standalone: no output unless --verbose is specified
- With --verbose: shows detailed logs with timestamps

This prevents the script from interfering with GNS3's logging
configuration while still allowing debug output when needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(web-wireshark): fix code quality issues from static analysis

Fix issues identified by flake8 and pylint static analysis:

Import and formatting fixes:
- Remove unused 'time' import
- Fix import order (stdlib before third-party)
- Remove unnecessary f-strings without placeholders
- Split long line to comply with flake8 (max line length: 127)

Code structure improvements:
- Remove unnecessary 'else' after 'return'
- Add 'from e' to exception re-raising for better tracebacks

Code quality metrics:
- pylint score: 8.11/10 → 9.97/10 (+1.86)
- flake8: 0 errors (max line length: 127, per CI/CD standard)
- All critical issues resolved

Note: R0914 (too-many-locals) warning is informational only,
code remains clear and maintainable.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): export LinkCapture schema in schemas module

Add LinkCapture to the schemas module exports in __init__.py to fix
the AttributeError when starting the GNS3 server.

Error was:
  AttributeError: module 'gns3server.schemas' has no attribute 'LinkCapture'

This was caused by adding the LinkCapture class to controller/links.py
but forgetting to export it in the schemas __init__.py.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(web-wireshark): add --image parameter for custom Docker images

Add optional --image parameter to allow using custom Docker images
for testing, instead of requiring gns3/web-wireshark:latest.

Changes:
- Add --image parameter to start command (default: gns3/web-wireshark:latest)
- Update get_or_create_container() to accept image parameter
- Update start_wireshark_session() to accept and pass image parameter
- Update TEST.md with examples of using custom images

Usage:
  # Using default image
  python3 manage_wireshark.py start --project-id x --link-id y --jwt-token z

  # Using custom image (for testing)
  python3 manage_wireshark.py start --project-id x --link-id y --jwt-token z --image ubuntu:latest

This makes it easier to test the script without having the official
gns3/web-wireshark image available.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(web-wireshark): add configurable resource limits and increase session capacity

Add configurable resource parameters and increase session support from 10 to 100.

Resource parameters:
- --memory: Memory limit (default: 2g)
- --memory-swap: Memory swap limit (default: same as memory)
- --cpus: CPU cores (default: 1.0)
- --pids-limit: Process limit (default: 1000)

Other improvements:
- Fix CPU quota calculation: 1000000 microseconds = 1.0 CPU core
  (was incorrectly 100000 = 0.1 CPU core)
- Add health check: xpra list command
- Add log configuration: json-file, max-size=10m, max-file=3
- Increase session capacity: 10 → 100 concurrent sessions
- Port range: 12300-12399 (was 12300-12309)
- Display range: :100-:199 (was :100-:109)

This matches the docker run command parameters and provides
better resource management and monitoring capabilities.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): update UDPLink.start_capture() to accept new parameters

Update UDPLink.start_capture() to accept the new wireshark and jwt_token
parameters and pass them to the parent Link class.

This fixes the TypeError when starting capture on UDP links:
  TypeError: UDPLink.start_capture() got an unexpected keyword argument 'wireshark'

The UDPLink class overrides start_capture() but didn't include the new
parameters added for Web Wireshark support.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs: add project memory infrastructure and JWT token flow documentation

- Add memory skill for recording project knowledge
- Document JWT token flow in Web Wireshark integration
- Update .gitignore to track skills and memory directories

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): fix Docker API URL format and use docker exec CLI

- Add "v" prefix to Docker API URL (http://docker/v{version}/...)
- Remove unused exec_create/exec_start API methods
- Fix Healthcheck.Test format to ["CMD-SHELL", "command"]
- Use docker exec CLI instead of Docker API for command execution
- This aligns with GNS3 docker_vm pattern

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): improve xpra session management and error handling

- Check and clean up existing sessions before starting new xpra session
- Add verification that xpra session started successfully
- Use docker exec CLI consistently instead of Docker API
- Improve error logging and diagnostics

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(web-wireshark): add timeout handling and container health checks

- Add REQUEST_TIMEOUT constant for Docker HTTP API requests
- Add asyncio.timeout wrapper for Docker API calls to prevent hanging
- Add _is_container_healthy() to check container responsiveness
- Add _exec_in_container() helper with timeout support
- Improve unhealthy container handling with force remove
- Add detailed logging for container health state

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(web-wireshark): split monolithic file into modules

Split manage_wireshark.py into:
- docker_client.py: Docker HTTP API client
- manager.py: WebWiresharkManager with session management
- manage_wireshark.py: CLI entry point

This improves code organization and maintainability.
Keep all timeout handling and health check logic.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): correct NanoCpus calculation to use nanoseconds

NanoCpus in Docker API requires nanoseconds (1 CPU = 1000000000),
not the previous incorrect multiplier of 100000.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): fix xpra command quoting and session verification

- Add quotes around --xvfb parameter value to preserve spaces
- Fix session verification to check display number instead of session name
  (xpra list shows "LIVE session at :185" not session name)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): run wireshark in background with &

Wireshark with curl pipeline runs continuously, so it must run
in background to avoid blocking. Also reduced timeout since we
don't wait for completion.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(web-wireshark): add deterministic hash for display/port allocation

Use MD5-based hash instead of Python's hash() which is randomized
across process restarts. This ensures display and port numbers are
stable when the GNS3 server restarts.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: extract link_id_to_port to shared utils module

Move deterministic hash functions to gns3server/utils/port_allocator.py
to avoid code duplication and ensure consistent algorithm across
manager.py and links.py.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(links): move import to top of file

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(web-wireshark): use correct Config access pattern like gns3-copilot

The Server config is an object with .protocol.value, .host, .port
attributes, not a dict. Fixes URL detection failing and falling back
to 127.0.0.1:3080 which doesn't work from inside containers.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add UUID validation utility and use in manage_wireshark

Create gns3server/utils/uuid_validator.py with validate_uuid function
to validate UUID format and catch typos early with clear error messages.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: use argparse.ArgumentTypeError for proper error message display

Previously ValueError was used which only showed 'invalid validate_uuid value'.
Now shows the full helpful message about expected format.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(web-wireshark): add --dpi=96 to xpra start parameters

Set standard DPI for better display scaling in web Wireshark.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: update TEST.md with correct project-id format

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(web-wireshark): hide xpra shutdown menu via XPRA_CLIENT_CAN_SHUTDOWN

Prevents users from accidentally shutting down the xpra server
through the HTML5 client menu.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(memory): add xpra-html5-client configuration reference

Document xpra HTML5 client configuration including:
- URL parameters for toolbar menu control
- default-settings.txt parameters (Features, Connection, Advanced)
- Server-controlled submenu items
- XPRA_CLIENT_CAN_SHUTDOWN environment variable
- Background image customization

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(web-wireshark): remove xpra stop cleanup to avoid zombie processes

xpra stop can leave zombie processes and timeout. Instead, let xpra start
reuse the display directly (it will overwrite existing session).

Also added --file-transfer=no, --printing=no, --sound=no to disable
unneeded features.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(web-wireshark): add fullscreen_button=false to default-settings.txt

Also remove --file-transfer/--printing/--sound from command line since
xpra doesn't support these options. Configure in default-settings.txt instead.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(web-wireshark): update xpra background with GNS3 branding

- Replace default xpra background with GNS3 icon and modern gradient
- Update CSS to use SVG background image with cover sizing
- Change background color to gradient from #021d3a to linear gradient
- Improves visual integration with GNS3 web interface

* style: add GPLv3 headers with copyright to web_wireshark modules

Co-Authored-By: YueGuobin <yueguobin@gmail.com>

* feat(docker): use Alibaba Cloud Debian mirror for faster builds in China

Use mirrors.aliyun.com for Debian packages with correct paths for
both main repo (/debian) and security repo (/debian-security).

Co-Authored-By: YueGuobin <yueguobin@gmail.com>

* feat(web-wireshark): update xpra background styling and echo command

- Replace `echo -e` with `printf` for better POSIX compatibility in Dockerfile
- Update xpra HTML5 client background to use GNS3 icon with light gradient
- Change background color from dark blue to light gray for improved visibility

* fix(web-wireshark): use pkill to stop xpra sessions

Replace 'xpra stop :{display}' with 'pkill -f "xpra.*:{display}"' for more reliable session termination.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(web-wireshark): start Wireshark in fullscreen mode

Add --fullscreen flag to Wireshark startup command to prevent window decorations and improve user experience.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): use carrier-grade NAT subnet to avoid conflicts

Change Docker network subnet from 172.28.0.0/16 to 100.64.1.0/24 to avoid conflicts with common private networks.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(web-wireshark): update subnet to /22 for 1000+ projects

Change Docker network subnet from 100.64.1.0/24 to 100.64.0.0/22 to support 1000+ projects (1022 available IPs). Update all documentation to reflect the new subnet.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): replace CGNAT subnet with configurable private subnet

Replace the CGNAT address block (100.64.0.0/22) with a standard private
subnet (172.31.0.0/22) to avoid network access issues. Many networks and
ISPs block or cannot route CGNAT addresses.

Changes:
- Add WebWiresharkSettings to config schema with configurable subnet
- Read network_subnet from config, default to 172.31.0.0/22
- Add Web Wireshark configuration section to sample config

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): comment out dpi setting

Allow xpra to use default DPI settings for better display scaling.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): set DPI in Xvfb to prevent scaling warnings

Add -dpi 96 to Xvfb startup command to match xpra's expected DPI
and avoid "scaling problems" warnings.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): remove custom Xvfb to use default Xorg-dummy

Remove --xvfb parameter to let xpra use the default Xorg-dummy driver,
which properly handles DPI changes when resize-display is enabled and
prevents DPI mismatch warnings.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* revert(web-wireshark): restore Xvfb configuration

Restore the Xvfb display server configuration for xpra sessions.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(web-wireshark): enhance xpra HTML5 default settings

Update xpra HTML5 configuration to disable additional unused features:
- audio, keyboard, mediasource, aurora, http-stream
- Improve readability using heredoc format instead of printf

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(api): add RBAC authentication to web wireshark WebSocket endpoint

- Add `has_privilege_on_websocket` dependency to enforce Link.Capture privilege
- Include current_user parameter in endpoint for user identification and logging
- Update endpoint documentation to reflect token requirement and privilege
- Add test example WebSocket URL in TEST.md for reference

* refactor(web-wireshark): add generic WebSocket proxy and disable xpra HTML

- Add websocket_to_websocket.py utility for binary WebSocket proxy
- Disable xpra HTML server with --html=no flag
- Update manager return value: url -> ws_url
- Simplify WebSocket proxy implementation in links.py

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(web-wireshark): remove unnecessary sleep delays

Remove 2-second sleep delays after container start and xpra initialization.
Health checks and verification happen immediately, improving startup time.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(agent): prevent blocking during project close

- Add 5s timeout for AgentService.checkpointer_conn.close()
- Optimize stop_all_sessions: single command instead of 100 docker exec calls
- Prevents indefinite hang when closing projects

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web-wireshark): return errors to client on startup failure

- Throw ControllerError when Web Wireshark startup fails
- Include detailed error messages from script stderr
- Support both ws_url and url response formats
- Client now receives proper error response instead of success

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(web-wireshark): parallelize startup and remove unnecessary waits

- Parallel execution: container info query + xpra start
- Remove xpra list verification (unnecessary)
- Fire-and-forget Wireshark startup (no timeout wait)
- Startup time reduced from ~7s to ~1s

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(web-wireshark): add stop-container and delete-container commands

- Add stop-container command to stop container on project close
- Add delete-container command to delete container on project delete
- Enables proper container lifecycle management
- Containers are now stopped (not deleted) when project closes
- Containers are deleted when project is deleted

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(project): simplify close flow by removing redundant xpra session cleanup

Since docker stop with timeout=0 already force-kills all processes in the
container, explicitly stopping xpra sessions before stopping the container
is unnecessary. Remove the _cleanup_web_wireshark_xpra_sessions() call
to reduce subprocess overhead and simplify the close flow.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* perf(docker): force kill containers on stop by default

Change docker stop timeout from 10 seconds to 0 (immediate SIGKILL).
Web Wireshark containers don't need graceful shutdown since they have
no persistent state.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(web-wireshark): enable WebSocket protocol for xpra connection

- Change xpra bind from --bind-tcp to --bind-ws for WebSocket support
- Enable HTML client with --html=on for xpra WebSocket server
- Pass binary subprotocol when connecting to xpra container

The xpra WebSocket server requires the 'binary' subprotocol during
handshake. Without it, the server returns 403 with message:
"client does not support 'binary' protocol".

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(web-wireshark): add WebSocket subprotocol negotiation support

Fix WebSocket proxy for xpra container by implementing proper subprotocol
negotiation. The xpra client requires the server to respond with the
negotiated subprotocol (binary) in the WebSocket handshake response.

Changes:
- Modified get_current_active_user_from_websocket to extract client's
  requested subprotocols from Sec-WebSocket-Protocol header
- Updated websocket.accept() call to include negotiated subprotocol
- Enhanced websocket_proxy to support subprotocol parameter
- Improved logging for WebSocket connection debugging

This fixes the issue where xpra clients would immediately disconnect
after connection due to missing subprotocol in server response.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(web-wireshark): remove verbose debug logging

Remove excessive debug logs from WebSocket proxy implementation while
keeping essential logs for troubleshooting:
- Keep: subprotocol negotiation, connection establishment, errors
- Remove: verbose step-by-step debugging information

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(web-wireshark): translate Chinese comments and docs to English

- Translate all Chinese comments in project.py, link.py, and links.py to English
- Convert TEST.md and maunal-test.md documentation to English
- Remove obsolete IMPLEMENTATION_PLAN.md

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore(docker): lock xpra version to 6.4.3 for reproducible builds

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* chore(docker): improve Dockerfile with TZ, no-install-recommends, and labels

- Add TZ=UTC timezone setting
- Use --no-install-recommends to reduce image size
- Add LABEL metadata for maintainer and description

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(docs): rename and consolidate documentation

- Rename TEST.md to WEB_WIRESHARK.md for clearer naming
- Merge maunal-test.md content into main document under "Manual Testing" section
- Remove obsolete gns3_icon_black.svg

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(docker): add xpra-x11 package and remove --no-install-recommends

- Add xpra-x11 package required for seamless mode
- Remove --no-install-recommends to ensure all recommended packages are installed

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(api): add capture file download endpoint

Add GET /{link_id}/capture/file endpoint to download PCAP capture files.
Supports downloading while capture is active (streaming).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(web_wireshark): improve session cleanup and add known issues

- Add cleanup of existing processes before starting xpra session to prevent
  "another window manager seems to be running" errors
- Stop all associated processes (xpra, wireshark, Xvfb) when stopping sessions
- Document known issues in WEB_WIRESHARK.md including JWT token security,
  duplicate code, and other implementation concerns

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(web_wireshark): add restart wireshark API endpoint

- Add POST /links/{link_id}/capture/wireshark/restart endpoint
- Add restart command to manage_wireshark.py CLI
- Add _restart_web_wireshark method in link.py controller
- Restart simply calls start_wireshark_session which handles cleanup

This allows users to recover after accidentally closing the Wireshark
window without having to stop and restart the entire capture.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: add Web Wireshark business process documentation

Document the Web Wireshark feature including architecture diagrams,
business processes for capture start/stop, container lifecycle,
WebSocket connection flow, and session management.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(web_wireshark): fix container network access and improve session cleanup

- Fix GNS3 server URL detection to use container gateway IP (172.31.0.1)
  instead of localhost/127.0.0.1/0.0.0.0 which don't work from containers
- Add import socket module for gateway IP conversion
- Move urllib.parse import to file header (code style improvement)
- Add logging for Web Wireshark startup (capture stream URL, display, etc.)
- Fix X lock file cleanup to prevent "Server is already active" errors
- Use exec pkill to reduce zombie processes from docker exec bash
- Add _cleanup_x_lock() method to remove X lock files after stopping sessions
- Update link.py to always log subprocess stderr for debugging

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web_wireshark): avoid zombie pkill processes by using Docker API

Replace pkill commands with Docker API-based process killing to prevent
zombie process accumulation in Web Wireshark containers.

Changes:
- Add DockerHTTPClient.list_processes() method to query container processes
- Rewrite _kill_process_tree() to use Docker API instead of pkill
- Unify all process cleanup to use _kill_process_tree() method
- Add logging for killed processes (count and PIDs)
- Include fallback to pkill if Docker API fails

This prevents the accumulation of zombie pkill processes that occurred
when using bash -c "pkill -9 -f pattern" commands.

Testing: Started and stopped multiple capture sessions, verified no new
pkill zombie processes are created during cleanup.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web_wireshark): use docker-init and cleanup socket files

Enable Docker init system (tini) as PID 1 and clean up xpra socket
files when stopping sessions to prevent zombie processes and leftover
sockets.

Changes:
- Add "Init": True to container host_config to use docker-init (tini)
- Extend _cleanup_x_lock() to remove xpra socket files
  * /run/user/1000/xpra/{display}/socket
  * /run/user/1000/xpra/*-{display}
  * /home/gns3/.xpra/*-{display}

Benefits:
- docker-init (tini) automatically reaps orphan processes, eliminating
  zombie process accumulation (tested: 53 zombies -> 0 zombies)
- Socket cleanup prevents xpra list from showing UNKNOWN sessions
- Cleaner container state after stopping sessions

Testing:
- Started and stopped 3 capture sessions
- Verified 0 zombie processes after stopping
- Verified xpra list shows "No xpra sessions found" (no UNKNOWN)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(web_wireshark): add comprehensive CLI script documentation

Add detailed module docstring to manage_wireshark.py explaining:
- Script purpose: CLI tool for manual management, debugging, and testing
- Important note: Use WebWiresharkManager directly for programmatic access
- Usage examples for all commands (start, stop, restart, stop-all, delete)
- Available commands list with descriptions
- Output format specification (JSON to stdout/stderr)
- Help information for getting command-specific usage

This clarifies that the script is intended as a CLI utility, not for
subprocess calls from within GNS3 server code.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(web_wireshark): use direct API calls instead of subprocess

Replace subprocess calls to manage_wireshark.py with direct API calls
to WebWiresharkManager, eliminating subprocess overhead and JSON parsing.

Changes:
- Add import for WebWiresharkManager
- Refactor _start_web_wireshark() to use direct API call
- Refactor _stop_web_wireshark() to use direct API call
- Refactor _restart_web_wireshark() to use direct API call

Benefits:
- Code reduction: 78 lines (-68%)
- Better performance: No subprocess creation overhead
- Better logging: Manager logs directly to GNS3 logging system
- Simpler error handling: Direct exceptions instead of return codes
- No JSON parsing: Direct Python objects
- Resource cleanup: Added finally blocks to ensure manager.close()

Testing:
- All Web Wireshark operations work correctly
- Logs now appear in GNS3 server logs instead of stderr
- Error handling improved with proper exception propagation

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web_wireshark): resolve circular import with delayed imports

Fix circular import error by moving WebWiresharkManager import from
module level to function level (delayed import).

Circular dependency was:
  link.py → WebWiresharkManager → Controller → link.py

Solution:
  - Remove top-level import of WebWiresharkManager
  - Add delayed imports in each method that uses it:
    * _start_web_wireshark()
    * _stop_web_wireshark()
    * _restart_web_wireshark()

This allows the controller module to fully initialize before importing
WebWiresharkManager, breaking the circular dependency.

Error before fix:
  ImportError: cannot import name 'Controller' from partially initialized
  module 'gns3server.controller' (most likely due to a circular import)

Testing:
  - Successfully imports Link module
  - GNS3 server starts without import errors

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor: move imports to module level and remove Controller dependency

Remove circular dependency between link.py and manager.py by:

1. Remove Controller import from manager.py
   - Manager no longer imports Controller
   - URL detection now relies on Config or default only
   - link.py passes capture_stream_url to manager

2. Move WebWiresharkManager import to module level in link.py
   - No longer need delayed imports since no circular dependency
   - Clean, standard Python import pattern

3. Remove redundant Config import in ensure_network()
   - Config was already imported at module level

Changes:
- manager.py: -24 lines (removed _get_gns3_url_from_controller and redundant import)
- link.py: -30 lines net reduction after moving import to top

Testing:
- All imports work without circular dependency errors
- GNS3 server starts successfully

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(docker_client): use text response for list_processes API

Docker API /containers/{id}/top returns plain text, not JSON.
The generic _request method auto-parses JSON which caused:
  "'dict' object has no attribute 'strip'"

Fix by using session.get() directly with response.text()
instead of the generic _request method.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web_wireshark): use correct capture stream URL endpoint

Remove passing capture_stream_url from link.py to manager.py.
Manager.py auto-detects URL using correct endpoint:
  /v3/projects/{project_id}/links/{link_id}/capture/stream

This endpoint supports JWT authentication, unlike the compute URL
(/v3/compute/...) which requires compute credentials.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web_wireshark): use container-local PIDs for process killing

- Docker API returns host PIDs, not container PIDs - use pgrep inside
  container to get correct PIDs for kill
- Fix list_processes to parse Docker API JSON response correctly
- Remove unnecessary exec prefix in shell commands
- Add logging for debugging process killing

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* perf(web_wireshark): optimize process killing with single pgrep call

Replace sequential _kill_process_tree calls with a new _kill_process_tree_batch
method that combines all patterns into a single regex. This reduces docker exec
calls from 8 to 1, improving performance from ~8 seconds to ~0.9 seconds.

Changes:
- Add _kill_process_tree_batch() method with combined regex pattern matching
- Update start_wireshark_session() to use batch cleanup
- Update stop_wireshark_session() to use batch cleanup
- Update stop_all_sessions() to use batch cleanup

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(web_wireshark): disable HTML5 client to improve xpra startup speed

Change --html=on to --html=off to disable the xpra HTML5 client interface.
This reduces xpra startup time from ~6.4s to ~2.9s (55% improvement) while
keeping the WebSocket server functional for browser connections.

The HTML5 client is not needed as we only require the WebSocket endpoint
for remote display forwarding.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(web_wireshark): optimize container health check using Docker built-in status

Replace manual docker exec ping with Docker's built-in health check status
to reduce container startup latency by ~1 second.

Changes:
- Use container["Health"]["Status"] instead of manual ping check
- Skip health check for containers with "healthy" or "starting" status
- Only verify containers with "unhealthy" status
- Keep manual health check for newly started containers

This reduces the container verification time from ~1s to near-zero
for healthy running containers while maintaining safety for unhealthy ones.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(web_wireshark): get gateway IP from Docker API instead of container exec

Replace slow docker exec method with fast Docker Network API call to get
gateway IP. This reduces gateway detection from ~850ms to ~1ms (1000x faster).

Changes:
- Use docker.get_network() API to get gateway from IPAM config
- Keep container exec methods as fallback if Docker API fails
- Remove container_id requirement (not needed for Docker API method)

The Docker network gateway is shared by all containers in the network,
so querying the network is more efficient than querying each container.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(web_wireshark): remove unreliable nameserver fallback for gateway detection

Remove the fallback method that reads /etc/resolv.conf nameserver as gateway.
This method is unreliable because:
- nameserver is not necessarily the gateway (could be upstream DNS)
- Many systems use 127.0.0.53 (systemd-resolved) or 127.0.0.1
- Even when not local, it could be public DNS (8.8.8.8, 1.1.1.1)

The current fallback strategy is sufficient:
1. Docker Network API (fast, reliable)
2. /proc/net/route (standard gateway detection)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(web_wireshark): use host perspective for process killing (37x faster)

Replace docker exec with host perspective process management for killing
container processes. This reduces process cleanup from ~850ms to ~23ms.

Key optimization:
- Get container init PID using docker inspect (fast)
- Use pgrep -P <init_pid> to find child processes from host perspective
- Kill processes directly using host PID (no docker exec needed)

Performance improvements:
- Process finding: 850ms → 23ms (37x faster)
- Process killing: 850ms → <1ms (850x faster)
- File checking: 850ms → 0.003ms (280,000x faster)

Fallback to docker exec method if host perspective fails.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(web_wireshark): use host perspective for file cleanup (891x faster)

Replace docker exec with direct filesystem access from host perspective
for cleaning up X lock files and xpra sockets. This reduces cleanup time
from ~890ms to ~1ms (891x faster).

Key optimization:
- Access /proc/<container_pid>/root/ directly from host
- Use os.remove() and glob.glob() instead of docker exec rm -f
- Fallback to docker exec if host perspective fails

This complements the earlier optimization for process killing, further
reducing startup time.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* Revert "perf(web_wireshark): use host perspective for file cleanup (891x faster)"

* fix(web_wireshark): recursively kill all descendant processes to prevent orphans

Previous implementation only killed direct children of container init using
`pgrep -P <init_pid>`, but xpra spawns child processes (Xvfb, pulseaudio,
ibus-daemon) that become grandchildren and were not being terminated.

This fix walks the entire process tree recursively to find and kill all
descendant processes, ensuring complete cleanup without orphans.

Performance: ~20-40ms (only 10-20ms slower than previous 23ms, but ensures
thorough cleanup).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(web_wireshark): add comprehensive startup/shutdown performance metrics

Add detailed performance characteristics section documenting:
- Startup performance breakdown (before/after optimization)
- Shutdown performance breakdown (before/after optimization)
- First startup vs subsequent startup comparison
- Measured production data
- Key optimization techniques

Performance improvements: 67% faster startup, 78% faster shutdown,
zero orphan processes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web_wireshark): ensure container stops on project close

Remove dependency on _web_wireshark_container_created flag which gets
reset when project is reloaded, causing container to not stop on project close.

Now directly attempts to stop container on every project close. If container
doesn't exist, the script handles it gracefully. This is simpler and more
reliable than maintaining a flag.

Also removed unnecessary _cleanup_web_wireshark_xpra_sessions() call since
stopping the container automatically terminates all xpra sessions inside.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web_wireshark): ensure container deletion on project delete

Remove dependency on _web_wireshark_container_created flag from
_cleanup_web_wireshark_container() to ensure container is deleted
when project is deleted, even if project was reloaded and flag was reset.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(web_wireshark): remove unused _web_wireshark_container_created flag

This flag was set but never read for any conditionals since we changed
the stop/delete methods to directly attempt the operation. It's purely
redundant code that adds confusion.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(web_wireshark): skip unnecessary cleanup using fast host perspective check

Add _check_residuals_exist() function that uses host perspective (~20ms)
to check if residual processes or socket files exist before cleanup.

Returns (has_process_residuals, has_socket_residuals) tuple to allow
selective cleanup - if only sockets need cleaning, skip the slower
process tree killing.

For new containers with no residuals, this saves ~3 seconds of
unnecessary docker exec calls.

Socket cleanup always uses docker exec for safety (avoiding accidental
host filesystem deletion from /proc/<pid>/root/ path errors).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(web_wireshark): increase container exec timeout from 5s to 10s

xpra start with Xvfb initialization can take longer than 5 seconds in Docker containers, causing timeout failures. Increasing timeout to 10 seconds to allow sufficient startup time.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(web_wireshark): optimize cleanup and document docker exec limitations

- Combine X lock and xpra socket cleanup into single docker exec call
  Reduces exec calls from 2 to 1 per stop operation, improving performance
  when stopping multiple capture sessions.

- Document Docker exec performance limitations with test data showing
  parallel exec is 42% slower than serial due to Docker daemon's
  internal queuing mechanism.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(links): prevent AssertionError in stream_pcap during rapid stop_capture

Fix race condition where stream_pcap checks link.capturing (True) but
link.capture_node becomes None before accessing link.compute.

This happens when stop_capture() is called and rapidly cleans up
_capture_node while WebSocket requests are still processing.

Now check both capturing and capture_node to ensure capture is still active.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(links): prevent AssertionError in stream_pcap during rapid stop_capture

Fix race condition where stream_pcap checks link.capturing (True) but
link.capture_node becomes None before accessing link.compute.

This happens when stop_capture() is called and rapidly cleans up
_capture_node while WebSocket requests are still processing.

Now check both capturing and capture_node to ensure capture is still active,
and log at DEBUG level since this is expected behavior during rapid stop.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(web_wireshark): load container config from settings

- Load memory, cpus, and pids_limit from WebWireshark config section
- Apply config only on container creation (_start_web_wireshark)
- Skip config on restart (_restart_web_wireshark) as container exists

Users can now configure container resources in gns3_server.conf:

[WebWireshark]
memory = 4g
cpus = 2.0
pids_limit = 2000

If not configured, defaults (2g, 1.0, 1000) are used.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(web_wireshark): add container statistics to /v3/statistics endpoint

Add Web Wireshark container monitoring to the statistics API for
better visibility into packet capture sessions.

Changes:
- Create new stats.py module with collect_webwireshark_stats()
  * Collects container info (status, project, active sessions)
  * Gets resource limits (memory, CPU, PIDs) from container config
  * Gets resource usage via docker stats (memory, CPU, PIDs)
  * Properly closes aiohttp connections to avoid leaks

- Update /v3/statistics endpoint to include webwireshark data
  * total_containers: Total number of Web Wireshark containers
  * running_containers: Number of containers currently running
  * active_sessions: Total active capture sessions
  * containers: Array with per-container details

- Update statistics-api.md documentation
  * Change URL from /v1/statistics to /v3/statistics
  * Add webwireshark field descriptions and examples
  * Add dashboard integration for Web Wireshark monitoring

Example response:
{
  "webwireshark": {
    "total_containers": 1,
    "running_containers": 1,
    "active_sessions": 2,
    "containers": [{
      "project_id": "...",
      "project_name": "test",
      "container_id": "6edc9029bac0",
      "status": "running",
      "memory_limit": "4.0 GB",
      "cpu_limit": "4.0",
      "pids_limit": 4000,
      "memory": "535.7MiB / 4GiB",
      "cpu": "0.29%",
      "pids": 124
    }]
  }
}

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(project): use WebWiresharkManager directly instead of subprocess

Replace subprocess calls to manage_wireshark.py with direct
WebWiresharkManager method calls in project.py. This eliminates
unnecessary process overhead and maintains consistent usage pattern
with link.py.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(logging): use %-formatting instead of f-strings in WebWireshark methods

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(web_wireshark): use info level log when container not found

Change log level from warning to info when container doesn't exist
during stop/delete operations, and fix f-string without placeholders.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(logging): lower log level for client disconnection

Change log level from warning to debug when client possibly
disconnects, as this is normal behavior when users close
browser tabs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(linting): resolve ruff warnings and errors

- Remove unused variable proc (fire-and-forget subprocess)
- Remove duplicate nodes property definition
- Remove extraneous f-prefix from f-strings without placeholders
- Change bare except to except Exception
- Remove unused imports (sys, asyncio, json)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(web_wireshark): improve error message when Docker image not found

When creating a container fails due to missing image, provide helpful
error message with docker pull command and local build instructions.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(schema): add wireshark field to Link schema

Add wireshark boolean field to Link Pydantic schema so it is
included in API responses when Web Wireshark session is active.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(link): add wireshark state tracking

Track whether Web Wireshark is running on a link by adding
_wireshark boolean property, updated on start/stop operations.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(chat): add session abort functionality

Add ability to abort streaming chat sessions via REST API.

Changes:
- Add abort flag to MessagesState for tracking abort requests
- Add session_id to state for abort event correlation
- Add _abort_flags dict and check/set/clear functions in gns3_copilot.py
- Add abort_handler_node to generate aborted tool messages
- Modify should_continue to route to abort_handler_node when aborting
- Add abort_session() method in AgentService
- Add POST /sessions/{session_id}/abort API endpoint
- Clear abort flag at stream start

When abort is triggered during streaming:
- If LLM has tool_calls pending, abort_handler_node generates
  aborted tool messages to maintain checkpoint consistency
- Prevents "insufficient tool messages" error on resume

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(chat): send tool_end event when stream is aborted

When a stream is aborted during tool execution, yield proper tool_end
events for aborted tools instead of a generic abort event. This maintains
compatibility with the frontend's expected tool_start/tool_end flow.

Changes:
- Add stream_aborted tracking flag
- After stream loop, check abort flag and yield tool_end events
- Add "abort" to ChatResponse type enum (unused but available)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(web_wireshark): add gns3-wireshark-setup command for Docker image setup

Add a new entry point script that allows users to setup the Web
Wireshark Docker image with a single command:

    pip install gns3-server && gns3-wireshark-setup

The script will:
1. Try to pull gns3/web-wireshark:latest from Docker Hub
2. If pull fails, build the image locally using the included Dockerfile
3. Show raw docker pull/build output for full visibility

Files changed:
- Add setup_wireshark_image.py (new entry point script)
- Update pyproject.toml (add gns3-wireshark-setup entry point)
- Update documentation (README.md, WEB_WIRESHARK.md, web-wireshark-business-process.md)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(appliance): add tags field support for appliances and templates

- Add tags field to ApplianceV1_6 and ApplianceV8 schemas
- Add tags propagation from appliance config to template
- Simplify VPCS builtin template tags

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* Fix test in test_link.py

* Fix link.py after failed tests

* feat(copilot): add device skills system for LLM context injection

- Add skills module with registry and DeviceSkillsTool
- Skills organized by vendor/series for easy extensibility
- Support device_type, category (device/protocol/feature), operation (config/diagnosis)
- First implementation: VPCS skill (gns3_vpcs_telnet)
- Skills tool integrated into teaching_assistant and lab_automation modes

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(vpcs): change prompt detection log from warning to debug

The "VPCS prompt not clearly detected" message is not an error,
connection works fine. Change to debug level to reduce noise.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(memory): add uBridge permission issue documentation

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(copilot): include link_id in links_summary output

Return link_id in the links_summary method so that the AI can
identify which link to analyze when using the packet capture tool.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(copilot): add packet capture analysis tool

Add PacketCaptureTool that allows AI to analyze packets from an active
GNS3 capture. The tool downloads capture files from the GNS3 server
and runs tshark analysis to help users understand network traffic.

Features:
- Download capture file from /capture/file endpoint
- Run tshark with custom arguments for flexible analysis
- Support analyzing specific packets (e.g., "explain packet #42")
- Support protocol statistics, traffic analysis, and expert info

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(copilot): use shlex.split to properly parse quoted tshark arguments

tshark_args may contain quoted filter expressions like "-Y \"frame.number == 24\"".
Using plain str.split() breaks these quoted arguments, causing tshark warnings
about conflicting display filters. Use shlex.split() to correctly parse quoted
arguments.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(copilot): simplify PacketCaptureTool to accept only packet_number

Instead of exposing complex tshark arguments to the LLM, the tool now
accepts a simple packet_number parameter and internally constructs the
tshark command with verbose output. This prevents the LLM from generating
incorrect tshark parameters while still providing detailed packet analysis.

Changes:
- Remove tshark_args and max_lines parameters
- Accept only packet_number as input
- Internal command: tshark -r <file> -Y "frame.number == N" -V
- Returns complete packet structure without line limits

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs: add CC BY-SA 4.0 license headers to documentation files

Add SPDX license headers referencing docs/LICENSE file to all markdown
and HTML documentation files in the docs directory.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(copilot): add topology planner skill

Add new skill for automatic network lab topology planning:
- IOU as default image, 10.0.0.0/8 IP range, max 10 nodes
- Node naming convention: R/S/PC + number
- 7-step workflow for create/rename/link/start/verify/config
- IP allocation rules and troubleshooting guide

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(copilot): add node positioning rules to topology planner skill

Add grid-based positioning for GNS3 nodes:
- Minimum 250px distance between nodes
- Grid layout: 4 columns, 300x250px spacing
- Formula: x = -400 + col * 300, y = -200 + row * 250
- Update workflow params_required to include x, y coordinates

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(copilot): remove unused helper functions from topology planner skill

Delete Python helper functions that cannot be serialized to JSON:
- calculate_node_positions()
- get_position_for_node()
- allocate_subnet()
- allocate_ip()
- TOPOLOGY_CREATION_STEPS (duplicate of workflow in skill dict)

LLM should use formula descriptions in skill to calculate positions/IPs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(copilot): add topology_planner workflow to lab automation prompt

Add device_skills tool to AVAILABLE TOOLS table.
Add TOPOLOGY PLANNING WORKFLOW section explaining:
- How to query topology_planner skill
- Default conventions (IOU, 10.0.0.0/8, naming rules, position formula)
- 7-step workflow for topology creation

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(copilot): support name parameter in gns3_create_node tool

- Add optional 'name' field to node creation, allowing direct naming
- Update Node constructor to pass name parameter
- Remove separate rename step (step_3) from topology planner workflow
- Renumber workflow steps: 1-2-3-4-5-6 (skipping old step_3 rename)

This eliminates the need for a separate gns3_update_node_name_tool call
when creating nodes with predefined names.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(copilot): update topology workflow to 6 steps

Remove rename step since name can be set directly in create_gns3_node.
Workflow changed from 7 steps to 6 steps.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(copilot): replace grid positioning with topology-based layout

Redesign node positioning rules to follow classic network topologies:
- Star: hub at center, spokes radiating outward
- Ring: nodes in circular arrangement
- Bus: linear chain of nodes
- Mesh: grid pattern for interconnected nodes
- Hierarchical: three-tier (Core → Distribution → Access)
- Linear P2P: point-to-point WAN links in a line

Remove old grid formula. Update prompt and workflow to reflect new approach.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* docs(copilot): add note about combining topology types

LLM can mix topology types as needed, e.g., star + linear_p2p for WAN segments.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(copilot): remove language matching rules from prompts

Remove language matching rules that cause inconsistent responses:
- "User writes in Chinese → Respond in Chinese"
- "User writes in English → Respond in English"

These created ambiguity and led to inconsistent language behavior.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(copilot): remove Chinese from response template

Replace mixed Chinese/English headers with English only:
- "操作总结 / Operation Summary" → "Operation Summary"
- "详细信息 / Details" → "Details"

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(copilot): remove Chinese mentions from prompts

Remove references to "Chinese" language in title_prompt.py:
- "Generates concise Chinese or English titles" → "Generates concise titles"
- "If the content is predominantly in Chinese, generate a Chinese title" → removed

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(copilot): use short_name for link labels and support short port names

- Match ports by both name and short_name for flexibility
- Use short_name (e.g. "e0/0") instead of full name (e.g. "Ethernet0/0")
  for link labels, improving visual clarity in dense topologies

* feat(topology): replace hardcoded interface names with dynamic placeholders

- Update topology planner skill template to use {short_name} placeholder
- Allows dynamic interface name generation based on device type
- Improves template flexibility for different network device configurations

* refactor(topology): use node-pair IP format and hyphenated naming

- Use 10.0.{node_pair}.x format for P2P links (e.g., 10.0.12.x for R-1-R-2)
- Change node naming to hyphenated format: R-1, R-2, SW-1, PC-1
- Use {short_name} as placeholder for port names in output template
- Remove unused IP_SUBNET_POOL constant
- Simplify ip_planning rules to intuitive format description

* fix(web_wireshark): raise minimum Docker API version to 1.44

Docker daemon requires API version 1.44+, so the minimum
supported version should match rather than defaulting to 1.40.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(web_wireshark): call ensure_network in start_wireshark_session

Ensure Docker network exists before creating container, so that
Web Wireshark works when started via API without going through CLI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(agent): improve Docker image setup with network error detection

- Add `is_network_error()` function to detect network-related failures in Docker pull output
- Modify `pull_image()` to capture output and return success status along with output
- Enhance main logic to detect network errors and skip local build when Docker Hub is inaccessible
- Provide helpful suggestions for network issues (Docker mirror, VPN, manual image transfer)
- Improve error messages to differentiate between network failures and other errors

* feat(docs): add Ubuntu 24.04 development setup guide

Add a comprehensive development environment setup guide for Ubuntu 24.04. The guide includes steps to install dependencies via PPA, configure Docker with mirror accelerators for users in mainland China, set up user permissions, and run the server from source with a Python virtual environment. This provides a clear, step-by-step reference for new contributors and developers.

* feat(docs): add pip mirror note for China mainland users

Add a comment in the development setup documentation suggesting the use of the Aliyun PyPI mirror for users in China mainland to improve installation speed and reliability. This helps overcome network restrictions and slow downloads from the default PyPI repository.

* feat(docs): add LVM root partition expansion guide

Add optional section to development setup documentation with instructions for expanding the root partition when using LVM. This helps developers resolve low disk space issues by utilizing unallocated space in the volume group. The guide includes commands to check LVM status, extend the logical volume, and verify the changes.

* feat(docs): add tshark to development setup dependencies

Add tshark to the apt install command in the development setup documentation. Tshark is required for packet capture functionality in GNS3, ensuring the development environment has all necessary tools for network analysis and debugging.

* feat: add GNS3 documentation skill with standardized structure

Add new documentation skill file defining standards for GNS3 technical documentation. The skill establishes a structured approach focusing on architecture diagrams, flow diagrams, and text descriptions while prohibiting code examples and user scenarios. This ensures consistent, high-quality documentation across the GNS3 ecosystem.

* feat(docs): update documentation skill to focus on server technical docs

Update the GNS3 documentation skill to specifically target server technical documentation under the docs/ directory. The revised standard emphasizes high-level understanding over implementation details, using ASCII diagrams for architecture and business processes, API endpoint tables, and measured performance data. Code specifics are intentionally omitted, directing readers to the codebase for implementation details.

* feat(docs): rewrite documentation skill and statistics API doc

- Rewrite gns3-documentation skill to match actual server-side doc style:
  focus on architecture/flow diagrams, API data, skip code details
- Switch diagram standard from ASCII to Mermaid (GitHub native rendering)
- Rewrite statistics-api.md with Mermaid architecture and sequence diagrams,
  add missing webwireshark container fields (memory_limit, cpu_limit, pids_limit),
  remove speculative content

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(docs): update VNC WebSocket console documentation with Mermaid diagrams

- Replace ASCII diagrams with Mermaid flowcharts for better visualization
- Clarify connection flow between browser, controller, compute, and VNC server
- Update authentication details and endpoint descriptions
- Add missing documentation for packet capture workflow
- Improve overall readability and maintainability

* docs: add AI disclaimer to documentation files

Add a standardized disclaimer to multiple documentation files indicating that the content has been organized by AI with reference to actual code. The disclaimer warns users that AI can make mistakes and advises verification against the source code when in doubt. This improves transparency about the documentation's origin and encourages careful usage.

* feat(docs): update chat API documentation with new features and clarifications

- Update title from "Design Document" to reflect current implementation status
- Add new features: session abort, copilot modes (teaching_assistant, lab_automation_assistant), session pinning
- Enhance architecture diagram with LangGraph StateGraph details including abort_handler_node, title_generator_node, and conditional edges
- Clarify statistics tracking: add ai_response_counted flag, filter title_generator_node from LLM counts, explain incremental token counting
- Improve documentation accuracy for real-time statistics collection and token calculation methods

* feat(docs): restructure command security documentation with visual workflows

- Replace verbose implementation details with concise architecture overview
- Add Mermaid diagrams to visualize command filtering and multi-line expansion flows
- Simplify configuration instructions and remove redundant examples
- Consolidate file structure and function reference tables for clarity
- Maintain all security principles while improving readability and maintainability

* feat(context): refactor context window management documentation for clarity

- Reorganize documentation with improved structure and visual diagrams
- Add Mermaid diagrams to illustrate architecture and trimming process
- Simplify content while maintaining technical accuracy
- Update token counting and trimming strategy explanations
- Enhance readability with better formatting and component tables

* feat(docs): update LLM model configs documentation

- Change config column type from JSONB to JSON (JSONB on PostgreSQL)
- Update provider table to mark base_url as required with planned optional status
- Add note about base_url being currently required in API schema
- Clarify GET own configs endpoint returns plain array
- Add copilot_mode field to update request schema
- Document update limitations for context_strategy and copilot_mode fields

* feat(docs): add GNS3StartNodeQuickTool and update node creation examples

- Add documentation for new `GNS3StartNodeQuickTool` (`start_gns3_node_quick`) that starts nodes without waiting for boot completion
- Update node creation example to include optional `name` field in template placement
- Clarify dynamic wait time strategy for `GNS3StartNodeTool` and contrast with quick tool
- Update tool file structure to reflect new configuration, display, and packet capture tools
- Fix truncated line in suspend tool documentation

* feat(docs): add comprehensive documentation structure with CC BY-SA 4.0 license

Add initial documentation directory with detailed README.md that outlines:
- Dual license structure (CC BY-SA 4.0 for docs, GPLv3 for code)
- Complete directory structure for features, AI Copilot, and bugs
- Feature documentation covering Controller+Compute setup, Statistics API, VNC WebSocket console, and Web Wireshark
- AI Copilot implemented features including Chat API, LLM model configs, command security, and multi-vendor device support

This provides organized technical documentation for the GNS3 server project with proper licensing and feature coverage.

* docs: update multi-vendor device support documentation

- Simplify vendor support table by removing redundant platform column
- Add VPCS as a simulator in status column
- Improve VPCS driver diagram with Mermaid syntax and detailed components
- Document VPCS-specific Netmiko parameters (fast_cli, global_delay_factor)
- Update VPCS tool usage example with connection options structure
- Remove outdated Nornir configuration approaches, reference current implementation
- Add file reference for VPCS tools location
- Clean up documentation structure for better readability

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: YueGuobin <yueguobin@gmail.com>
Co-authored-by: Jeremy Grossmann <grossmj@gns3.net>
2026-04-20 22:56:55 +08:00

1554 lines
52 KiB
Python

#!/usr/bin/env python
#
# Copyright (C) 2016 GNS3 Technologies Inc.
#
# This program is free software: you can redistribute it and/or modify
# it under the terms of the GNU General Public License as published by
# the Free Software Foundation, either version 3 of the License, or
# (at your option) any later version.
#
# This program is distributed in the hope that it will be useful,
# but WITHOUT ANY WARRANTY; without even the implied warranty of
# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
# GNU General Public License for more details.
#
# You should have received a copy of the GNU General Public License
# along with this program. If not, see <http://www.gnu.org/licenses/>.
import re
import os
import json
import uuid
import copy
import shutil
import time
import asyncio
import aiofiles
import tempfile
import zipfile
import pathlib
from uuid import UUID, uuid4
from fastapi import HTTPException, status
from .node import Node
from .compute import ComputeError
from .snapshot import Snapshot
from .drawing import Drawing
from .topology import project_to_topology, load_topology
from .udp_link import UDPLink
from ..config import Config
from ..utils.path import check_path_allowed, get_default_project_directory
from ..utils.application_id import get_next_application_id
from ..utils.asyncio.pool import Pool
from ..utils.asyncio import locking
from ..utils.asyncio import aiozipstream
from ..utils.asyncio import wait_run_in_executor
from .export_project import export_project
from .import_project import import_project, update_snapshots, regenerate_topology_ids
from .controller_error import ControllerError, ControllerForbiddenError, ControllerNotFoundError
from gns3server.agent.web_wireshark.manager import WebWiresharkManager
import logging
log = logging.getLogger(__name__)
def open_required(func):
"""
Use this decorator to raise an error if the project is not opened
"""
def wrapper(self, *args, **kwargs):
if self._status == "closed":
raise ControllerForbiddenError("The project is not opened")
return func(self, *args, **kwargs)
return wrapper
class Project:
"""
A project inside a controller
:param project_id: force project identifier (None by default auto generate an UUID)
:param path: path of the project. (None use the standard directory)
:param status: Status of the project (opened / closed)
"""
def __init__(
self,
name=None,
project_id=None,
path=None,
controller=None,
status="opened",
filename=None,
auto_start=False,
auto_open=False,
auto_close=True,
scene_height=1000,
scene_width=2000,
zoom=100,
show_layers=False,
snap_to_grid=False,
show_grid=False,
grid_size=75,
drawing_grid_size=25,
show_interface_labels=False,
variables=None,
supplier=None,
created_by=None,
):
self._controller = controller
assert name is not None
self._name = name
self._auto_start = auto_start
self._auto_close = auto_close
self._auto_open = auto_open
self._status = status
self._scene_height = scene_height
self._scene_width = scene_width
self._zoom = zoom
self._show_layers = show_layers
self._snap_to_grid = snap_to_grid
self._show_grid = show_grid
self._grid_size = grid_size
self._drawing_grid_size = drawing_grid_size
self._show_interface_labels = show_interface_labels
self._variables = variables
self._supplier = supplier
self._created_by = created_by
self._snapshots_config_file = "snapshots.conf"
self._loading = False
self._closing = False
# Disallow overwrite of existing project
if project_id is None and path is not None:
if os.path.exists(path):
raise ControllerForbiddenError(f"The path {path} already exists")
else:
raise ControllerForbiddenError("Providing a path to create a new project is deprecated.")
if project_id is None:
self._id = str(uuid4())
else:
try:
UUID(project_id, version=4)
except ValueError:
raise ControllerError(f"{project_id} is not a valid UUID")
self._id = project_id
if path is None:
path = os.path.join(get_default_project_directory(), self._id)
self.path = path
if filename is not None:
self._filename = filename
else:
self._filename = self.name + ".gns3"
self.reset()
# At project creation we write an empty .gns3 with the meta
if not os.path.exists(self._topology_file()):
assert self._status != "closed"
self.dump()
self._iou_id_lock = asyncio.Lock()
log.debug(f'Project "{self.name}" [{self._id}] loaded')
self.emit_controller_notification("project.created", self.asdict())
def emit_notification(self, action, event):
"""
Emit a project notification to all clients using this project.
:param action: Action name
:param event: Event to send
"""
self._controller.notification.project_emit(action, event, project_id=self.id)
def emit_controller_notification(self, action, event):
"""
Emit a controller notification, all clients will see it.
:param action: Action name
:param event: Event to send
"""
self._controller.notification.controller_emit(action, event)
async def update(self, **kwargs):
"""
Update the project
:param kwargs: Project properties
"""
old_json = self.asdict()
for prop in kwargs:
setattr(self, prop, kwargs[prop])
# We send notif only if object has changed
if old_json != self.asdict():
self.emit_controller_notification("project.updated", self.asdict())
self.dump()
# update on computes
for compute in list(self._project_created_on_compute):
await compute.put(f"/projects/{self._id}", {"variables": self.variables})
def reset(self):
"""
Called when open/close a project. Cleanup internal stuff
"""
self._allocated_node_names = set()
self._nodes = {}
self._links = {}
self._drawings = {}
self._snapshots = {}
self._computes = []
self._load_snapshot_config()
# Create the project on demand on the compute node
self._project_created_on_compute = set()
@property
def scene_height(self):
return self._scene_height
@scene_height.setter
def scene_height(self, val):
"""
Height of the drawing area
"""
self._scene_height = val
@property
def scene_width(self):
return self._scene_width
@scene_width.setter
def scene_width(self, val):
"""
Width of the drawing area
"""
self._scene_width = val
@property
def zoom(self):
"""
Zoom level in percentage
:return: integer > 0
"""
return self._zoom
@zoom.setter
def zoom(self, zoom):
"""
Setter for zoom level in percentage
"""
self._zoom = zoom
@property
def show_layers(self):
"""
Show layers mode
:return: bool
"""
return self._show_layers
@show_layers.setter
def show_layers(self, show_layers):
"""
Setter for show layers mode
"""
self._show_layers = show_layers
@property
def snap_to_grid(self):
"""
Snap to grid mode
:return: bool
"""
return self._snap_to_grid
@snap_to_grid.setter
def snap_to_grid(self, snap_to_grid):
"""
Setter for snap to grid mode
"""
self._snap_to_grid = snap_to_grid
@property
def show_grid(self):
"""
Show grid mode
:return: bool
"""
return self._show_grid
@show_grid.setter
def show_grid(self, show_grid):
"""
Setter for showing the grid mode
"""
self._show_grid = show_grid
@property
def grid_size(self):
"""
Grid size
:return: integer
"""
return self._grid_size
@grid_size.setter
def grid_size(self, grid_size):
"""
Setter for grid size
"""
self._grid_size = grid_size
@property
def drawing_grid_size(self):
"""
Grid size
:return: integer
"""
return self._drawing_grid_size
@drawing_grid_size.setter
def drawing_grid_size(self, grid_size):
"""
Setter for grid size
"""
self._drawing_grid_size = grid_size
@property
def show_interface_labels(self):
"""
Show interface labels mode
:return: bool
"""
return self._show_interface_labels
@show_interface_labels.setter
def show_interface_labels(self, show_interface_labels):
"""
Setter for show interface labels
"""
self._show_interface_labels = show_interface_labels
@property
def variables(self):
"""
Variables applied to the project
:return: list
"""
return self._variables
@variables.setter
def variables(self, variables):
"""
Setter for variables applied to the project
"""
self._variables = variables
@property
def supplier(self):
"""
Supplier of the project
:return: dict
"""
return self._supplier
@supplier.setter
def supplier(self, supplier):
"""
Setter for supplier of the project
"""
self._supplier = supplier
@property
def created_by(self):
"""
Username of the user who created the project
:return: str or None
"""
return self._created_by
@created_by.setter
def created_by(self, created_by):
"""
Setter for the username of the user who created the project
"""
self._created_by = created_by
@property
def auto_start(self):
"""
Should project auto start when opened
"""
return self._auto_start
@auto_start.setter
def auto_start(self, val):
self._auto_start = val
@property
def auto_close(self):
"""
Should project automatically closed when client
stop listening for notification
"""
return self._auto_close
@auto_close.setter
def auto_close(self, val):
self._auto_close = val
@property
def auto_open(self):
return self._auto_open
@auto_open.setter
def auto_open(self, val):
self._auto_open = val
@property
def controller(self):
return self._controller
@property
def name(self):
return self._name
@name.setter
def name(self, val):
self._name = val
@property
def id(self):
return self._id
@property
def path(self):
return self._path
@property
def status(self):
return self._status
@path.setter
def path(self, path):
check_path_allowed(path)
try:
os.makedirs(path, exist_ok=True)
except OSError as e:
raise ControllerError(f"Could not create project directory: {e}")
if '"' in path:
raise ControllerForbiddenError(
'You are not allowed to use " in the project directory path. Not supported by Dynamips.'
)
self._path = path
@property
def captures_directory(self):
"""
Location of the captures files
"""
path = os.path.join(self._path, "project-files", "captures")
os.makedirs(path, exist_ok=True)
return path
@property
def pictures_directory(self):
"""
Location of the images files
"""
path = os.path.join(self._path, "project-files", "images")
os.makedirs(path, exist_ok=True)
return path
@property
def computes(self):
"""
:return: List of computes used by the project
"""
if self._status == "closed":
return self._get_closed_data("computes", "compute_id").values()
return self._project_created_on_compute
def remove_allocated_node_name(self, name):
"""
Removes an allocated node name
:param name: allocated node name
"""
if name in self._allocated_node_names:
self._allocated_node_names.remove(name)
def update_allocated_node_name(self, base_name):
"""
Updates a node name or generate a new if no node
name is available.
:param base_name: new node base name
"""
if base_name is None:
return None
base_name = re.sub(r"[ ]", "", base_name) # remove spaces in node name
if base_name in self._allocated_node_names:
base_name = re.sub(r"[0-9]+$", "{0}", base_name)
if "{0}" in base_name or "{id}" in base_name:
# base name is a template, replace {0} or {id} by an unique identifier
for number in range(1, 1000000):
try:
name = base_name.format(number, id=number, name="Node")
except KeyError as e:
raise ControllerError("{" + e.args[0] + "} is not a valid replacement string in the node name")
except (ValueError, IndexError):
raise ControllerError(f"{base_name} is not a valid replacement string in the node name")
if name not in self._allocated_node_names:
self._allocated_node_names.add(name)
return name
else:
if base_name not in self._allocated_node_names:
self._allocated_node_names.add(base_name)
return base_name
# base name is not unique, let's find a unique name by appending a number
for number in range(1, 1000000):
name = base_name + str(number)
if name not in self._allocated_node_names:
self._allocated_node_names.add(name)
return name
raise ControllerError("A node name could not be allocated (node limit reached?)")
def update_node_name(self, node, new_name):
if new_name and node.name != new_name:
self.remove_allocated_node_name(node.name)
return self.update_allocated_node_name(new_name)
return new_name
@open_required
async def add_node_from_template(self, template, x=0, y=0, name=None, compute_id=None):
"""
Create a node from a template.
"""
template["x"] = x
template["y"] = y
node_type = template.pop("template_type")
if compute_id:
# use a custom compute_id
compute = self.controller.get_compute(compute_id)
else:
compute = self.controller.get_compute(template.pop("compute_id"))
template_name = template.pop("name")
log.info(f'Creating node from template "{template_name}" on compute "{compute.name}" [{compute.id}]')
default_name_format = template.pop("default_name_format", "{name}-{0}")
if name is None:
name = default_name_format.replace("{name}", template_name)
node_id = str(uuid.uuid4())
node = await self.add_node(compute, name, node_id, node_type=node_type, **template)
return node
async def _create_node(self, compute, name, node_id, node_type=None, **kwargs):
node = Node(self, compute, name, node_id=node_id, node_type=node_type, **kwargs)
if compute not in self._project_created_on_compute:
# For a local server we send the project path
if compute.id == "local":
data = {"name": self._name, "project_id": self._id, "path": self._path}
else:
data = {"name": self._name, "project_id": self._id}
if self._variables:
data["variables"] = self._variables
await compute.post("/projects", data=data)
self._project_created_on_compute.add(compute)
await node.create()
self._nodes[node.id] = node
return node
@open_required
async def add_node(self, compute, name, node_id, dump=True, node_type=None, **kwargs):
"""
Create a node or return an existing node
:param dump: Dump topology to disk
:param kwargs: See the documentation of node
"""
if node_id in self._nodes:
return self._nodes[node_id]
if compute.id not in self._computes:
self._computes.append(compute.id)
if node_type == "iou":
async with self._iou_id_lock:
# wait for an IOU node to be completely created before adding a new one
# this is important otherwise we allocate the same application ID (used
# to generate MAC addresses) when creating multiple IOU node at the same time
if "properties" in kwargs.keys():
# allocate a new application id for nodes loaded from the project
kwargs.get("properties")["application_id"] = get_next_application_id(
self._controller.projects, self._computes
)
elif "application_id" not in kwargs.keys() and not kwargs.get("properties"):
# allocate a new application id for nodes added to the project
kwargs["application_id"] = get_next_application_id(self._controller.projects, self._computes)
node = await self._create_node(compute, name, node_id, node_type, **kwargs)
else:
node = await self._create_node(compute, name, node_id, node_type, **kwargs)
self.emit_notification("node.created", node.asdict())
if dump:
self.dump()
return node
@locking
async def __delete_node_links(self, node):
"""
Delete all link connected to this node.
The operation use a lock to avoid cleaning links from
multiple nodes at the same time.
"""
for link in list(self._links.values()):
if node in link.nodes:
await self.delete_link(link.id, force_delete=True)
@open_required
async def delete_node(self, node_id):
node = self.get_node(node_id)
if node.locked:
raise ControllerError(f"Node {node.name} cannot be deleted because it is locked")
await self.__delete_node_links(node)
self.remove_allocated_node_name(node.name)
del self._nodes[node.id]
await node.destroy()
# refresh the compute IDs list
self._computes = [n.compute.id for n in self.nodes.values()]
self.dump()
self.emit_notification("node.deleted", node.asdict())
@open_required
def get_node(self, node_id):
"""
Return the node or raise a 404 if the node is unknown
"""
try:
return self._nodes[node_id]
except KeyError:
raise ControllerNotFoundError(f"Node ID {node_id} doesn't exist")
def _get_closed_data(self, section, id_key):
"""
Get the data for a project from the .gns3 when
the project is closed
:param section: The section name in the .gns3
:param id_key: The key for the element unique id
"""
try:
path = self._topology_file()
with open(path) as f:
topology = json.load(f)
except OSError as e:
raise ControllerError(f"Could not load topology: {e}")
try:
data = {}
for elem in topology["topology"][section]:
data[elem[id_key]] = elem
return data
except KeyError:
raise ControllerNotFoundError(f"Section {section} not found in the topology")
@property
def nodes(self):
"""
:returns: Dictionary of the nodes
"""
if self._status == "closed":
return self._get_closed_data("nodes", "node_id")
return self._nodes
@property
def drawings(self):
"""
:returns: Dictionary of the drawings
"""
if self._status == "closed":
return self._get_closed_data("drawings", "drawing_id")
return self._drawings
@open_required
async def add_drawing(self, drawing_id=None, dump=True, **kwargs):
"""
Create an drawing or return an existing drawing
:param dump: Dump the topology to disk
:param kwargs: See the documentation of drawing
"""
if drawing_id not in self._drawings:
drawing = Drawing(self, drawing_id=drawing_id, **kwargs)
self._drawings[drawing.id] = drawing
self.emit_notification("drawing.created", drawing.asdict())
if dump:
self.dump()
return drawing
return self._drawings[drawing_id]
@open_required
def get_drawing(self, drawing_id):
"""
Return the Drawing or raise a 404 if the drawing is unknown
"""
try:
return self._drawings[drawing_id]
except KeyError:
raise ControllerNotFoundError(f"Drawing ID {drawing_id} doesn't exist")
@open_required
async def delete_drawing(self, drawing_id):
drawing = self.get_drawing(drawing_id)
if drawing.locked:
raise ControllerError(f"Drawing ID {drawing_id} cannot be deleted because it is locked")
del self._drawings[drawing.id]
self.dump()
self.emit_notification("drawing.deleted", drawing.asdict())
@open_required
async def add_link(self, link_id=None, dump=True):
"""
Create a link. By default the link is empty
:param dump: Dump topology to disk
"""
if link_id and link_id in self._links:
return self._links[link_id]
link = UDPLink(self, link_id=link_id)
self._links[link.id] = link
if dump:
self.dump()
return link
@open_required
async def delete_link(self, link_id, force_delete=False):
link = self.get_link(link_id)
del self._links[link.id]
try:
await link.delete()
except Exception:
if force_delete is False:
raise
self.dump()
self.emit_notification("link.deleted", link.asdict())
@open_required
def get_link(self, link_id):
"""
Return the Link or raise a 404 if the link is unknown
"""
try:
return self._links[link_id]
except KeyError:
raise ControllerNotFoundError(f"Link ID {link_id} doesn't exist")
@property
def links(self):
"""
:returns: Dictionary of the Links
"""
if self._status == "closed":
return self._get_closed_data("links", "link_id")
return self._links
@property
def snapshots(self):
"""
:returns: Dictionary of snapshots
"""
return self._snapshots
@open_required
def get_snapshot(self, snapshot_id):
"""
Return the snapshot or raise a 404 if the snapshot is unknown
"""
try:
return self._snapshots[snapshot_id]
except KeyError:
raise ControllerNotFoundError(f"Snapshot ID {snapshot_id} doesn't exist")
def _load_snapshot_config(self):
snapshot_dir = os.path.join(self.path, "snapshots")
self._snapshot_conf_path = os.path.join(snapshot_dir, self._snapshots_config_file)
self._snapshot_conf = []
if os.path.isfile(self._snapshot_conf_path):
try:
with open(self._snapshot_conf_path, encoding="utf-8") as f:
self._snapshot_conf = json.load(f)
except (OSError, UnicodeDecodeError, ValueError) as e:
raise ControllerError(f"Could not read snapshot config {e}")
# Load all legacy snapshots (.gns3project files) to create an initial snapshot config if it doesn't exist
if os.path.exists(snapshot_dir) and not self._snapshot_conf:
for snap in os.listdir(snapshot_dir):
if snap.endswith(".gns3project"):
try:
snapshot = Snapshot(self, filename=snap)
except ValueError:
log.error("Invalid snapshot file: {}".format(snap))
continue
self._snapshots[snapshot.id] = snapshot
else:
# Create the Snapshot instances from the snapshot config file
for snapshot_entry in self._snapshot_conf:
try:
path = os.path.join(snapshot_dir, snapshot_entry["filename"])
if not os.path.isfile(path):
log.warning("Snapshot file '{}' does not exist".format(path))
continue
snapshot_entry.pop("project_id")
snapshot = Snapshot(self, **snapshot_entry)
self._snapshots[snapshot.id] = snapshot
except KeyError:
log.error("Invalid entry in snapshot config file: {}".format(snapshot_entry))
continue
self._save_snapshot_config()
def _save_snapshot_config(self):
if not self._snapshots:
return
self._snapshot_conf = []
for snapshot in self._snapshots.values():
self._snapshot_conf.append(snapshot.asdict())
try:
with open(self._snapshot_conf_path, 'w+') as f:
json.dump(self._snapshot_conf, f, indent=4)
except OSError as e:
log.error("Cannot write snapshot config '{}': {}".format(self._snapshot_conf_path, e))
@open_required
async def snapshot(self, name):
"""
Snapshot the project
:param name: Name of the snapshot
"""
if name in [snap.name for snap in self._snapshots.values()]:
raise ControllerError(f"The snapshot name {name} already exists")
snapshot = Snapshot(self, name=name)
await snapshot.create()
self._snapshots[snapshot.id] = snapshot
self._save_snapshot_config()
return snapshot
@open_required
async def delete_snapshot(self, snapshot_id):
snapshot = self.get_snapshot(snapshot_id)
del self._snapshots[snapshot.id]
self._save_snapshot_config()
os.remove(snapshot.path)
@locking
async def close(self, ignore_notification=False):
if self._status == "closed" or self._closing:
return
if self._loading:
log.warning(f"Closing project '{self.name}' ignored because it is being loaded")
return
self._closing = True
try:
await self.stop_all()
except HTTPException as e:
if not e.status_code == status.HTTP_405_METHOD_NOT_ALLOWED:
raise
for compute in list(self._project_created_on_compute):
try:
await compute.post(f"/projects/{self._id}/close", dont_connect=True)
# We don't care if a compute is down at this step
except (ComputeError, ControllerError, TimeoutError):
pass
self._clean_pictures()
self._status = "closed"
if not ignore_notification:
self.emit_controller_notification("project.closed", self.asdict())
# Stop Web Wireshark container (all xpra sessions terminate with container)
await self._stop_web_wireshark_container()
# Cleanup GNS3 Copilot AgentService for this project
await self._cleanup_copilot_agent()
self.reset()
self._closing = False
def _clean_pictures(self):
"""
Delete unused pictures.
"""
# Project have been deleted or is loading or is not opened
if not os.path.exists(self.path) or self._loading or self._status != "opened":
return
try:
pictures = set(os.listdir(self.pictures_directory))
for drawing in self._drawings.values():
try:
resource_filename = drawing.resource_filename
if resource_filename:
pictures.remove(resource_filename)
except KeyError:
pass
# don't remove supplier's logo
if self.supplier:
try:
logo = self.supplier["logo"]
pictures.remove(logo)
except KeyError:
pass
for pic_filename in pictures:
path = os.path.join(self.pictures_directory, pic_filename)
log.info(f"Deleting unused picture '{path}'")
os.remove(path)
except OSError as e:
log.warning(f"Could not delete unused pictures: {e}")
async def _cleanup_web_wireshark_xpra_sessions(self):
"""
Cleanup all Web Wireshark xpra sessions (without deleting container).
Called when project is closed to stop all xpra sessions and Wireshark processes,
while keeping the container for quick reuse when project is reopened.
"""
try:
log.info("Stopping xpra sessions for project '%s' (%s)", self.name, self._id)
manager = WebWiresharkManager()
try:
await manager.stop_all_sessions(self._id)
log.info("Web Wireshark xpra sessions stopped successfully")
finally:
await manager.close()
except Exception as e:
# Don't raise exception to avoid affecting project close flow
log.warning("Failed to cleanup xpra sessions for project '%s': %s", self.name, e)
async def _stop_web_wireshark_container(self):
"""
Stop Web Wireshark container (without deleting).
Called when project is closed to stop the container and free memory,
while keeping the container for quick startup when project is reopened.
"""
try:
container_name = f"gns3-wireshark-{self._id}"
log.info("Stopping Web Wireshark container '%s' for project '%s'", container_name, self.name)
manager = WebWiresharkManager()
try:
await manager.stop_container(self._id)
log.info("Web Wireshark container stopped successfully")
finally:
await manager.close()
except Exception as e:
# Don't fail project close if container stop fails
log.warning("Failed to stop container for project '%s': %s", self.name, e)
async def _cleanup_web_wireshark_container(self):
"""
Delete Web Wireshark container.
Called when project is deleted to stop and remove the container.
"""
try:
container_name = f"gns3-wireshark-{self._id}"
log.info("Deleting Web Wireshark container '%s' for project '%s'", container_name, self.name)
manager = WebWiresharkManager()
try:
await manager.delete_container(self._id)
log.info("Web Wireshark container deleted successfully")
finally:
await manager.close()
except Exception as e:
# Don't fail project delete if container cleanup fails
log.warning("Failed to delete container for project '%s': %s", self.name, e)
async def _cleanup_copilot_agent(self):
"""
Cleanup GNS3 Copilot AgentService for this project.
This should be called when the project is closed to free resources.
"""
try:
from gns3server.agent.gns3_copilot.project_agent_manager import get_project_agent_manager
agent_manager = await get_project_agent_manager()
if agent_manager.has_agent(self._id):
log.info("Cleaning up AgentService for project '%s' (%s)", self.name, self._id)
await agent_manager.remove_agent(self._id)
except Exception as e:
# Don't fail project close if agent cleanup fails
log.warning("Failed to cleanup AgentService for project '%s': %s", self.name, e)
async def delete(self):
if self._status != "opened":
try:
await self.open()
except ControllerError as e:
# ignore missing images or other conflicts when deleting a project
log.warning(f"Conflict while deleting project: {e}")
await self.delete_on_computes()
await self.close()
# Delete Web Wireshark container
await self._cleanup_web_wireshark_container()
try:
project_directory = get_default_project_directory()
if not os.path.commonprefix([project_directory, self.path]) == project_directory:
raise ControllerError(
f"Project '{self._name}' cannot be deleted because it is not in the default project directory: '{project_directory}'"
)
shutil.rmtree(self.path)
except OSError as e:
raise ControllerError(f"Cannot delete project directory {self.path}: {str(e)}")
self.emit_controller_notification("project.deleted", self.asdict())
async def delete_on_computes(self):
"""
Delete the project on computes but not on controller
"""
for compute in list(self._project_created_on_compute):
if compute.id != "local":
await compute.delete(f"/projects/{self._id}")
self._project_created_on_compute.remove(compute)
@classmethod
def _get_default_project_directory(cls):
"""
Return the default location for the project directory
depending on the operating system
"""
server_config = Config.instance().settings.Server
path = os.path.expanduser(server_config.projects_path)
path = os.path.normpath(path)
try:
os.makedirs(path, exist_ok=True)
except OSError as e:
raise ControllerError(f"Could not create project directory: {e}")
return path
def _topology_file(self):
return os.path.join(self.path, self._filename)
@locking
async def open(self):
"""
Load topology elements
"""
if self._closing is True:
raise ControllerError("Project is closing, please try again in a few seconds...")
if self._status == "opened":
return
self.reset()
self._loading = True
self._status = "opened"
path = self._topology_file()
if not os.path.exists(path):
self._loading = False
return
try:
shutil.copy(path, path + ".backup")
except OSError:
pass
try:
project_data = load_topology(path)
# load meta of project
keys_to_load = [
"auto_start",
"auto_close",
"auto_open",
"scene_height",
"scene_width",
"zoom",
"show_layers",
"snap_to_grid",
"show_grid",
"grid_size",
"drawing_grid_size",
"show_interface_labels",
]
for key in keys_to_load:
val = project_data.get(key, None)
if val is not None:
setattr(self, key, val)
topology = project_data["topology"]
for compute in topology.get("computes", []):
await self.controller.add_compute(**compute)
# Get all compute used in the project
# used to allocate application IDs for IOU nodes.
for node in topology.get("nodes", []):
compute_id = node.get("compute_id")
if compute_id not in self._computes:
self._computes.append(compute_id)
for node in topology.get("nodes", []):
compute = self.controller.get_compute(node.pop("compute_id"))
name = node.pop("name")
node_id = node.pop("node_id", str(uuid.uuid4()))
await self.add_node(compute, name, node_id, dump=False, **node)
for link_data in topology.get("links", []):
if "link_id" not in link_data.keys():
# skip the link
continue
link = await self.add_link(link_id=link_data["link_id"])
if "filters" in link_data:
await link.update_filters(link_data["filters"])
if "link_style" in link_data:
await link.update_link_style(link_data["link_style"])
for node_link in link_data.get("nodes", []):
node = self.get_node(node_link["node_id"])
port = node.get_port(node_link["adapter_number"], node_link["port_number"])
if port is None:
log.warning(
"Port {}/{} for {} not found".format(
node_link["adapter_number"], node_link["port_number"], node.name
)
)
continue
if port.link is not None:
log.warning(
"Port {}/{} is already connected to link ID {}".format(
node_link["adapter_number"], node_link["port_number"], port.link.id
)
)
continue
await link.add_node(
node,
node_link["adapter_number"],
node_link["port_number"],
label=node_link.get("label"),
dump=False,
)
if len(link.nodes) != 2:
# a link should have 2 attached nodes, this can happen with corrupted projects
await self.delete_link(link.id, force_delete=True)
for drawing_data in topology.get("drawings", []):
await self.add_drawing(dump=False, **drawing_data)
self.dump()
# We catch all error to be able to roll back the .gns3 to the previous state
except Exception as e:
for compute in list(self._project_created_on_compute):
try:
await compute.post(f"/projects/{self._id}/close")
# We don't care if a compute is down at this step
except ComputeError:
pass
try:
if os.path.exists(path + ".backup"):
shutil.copy(path + ".backup", path)
except OSError:
pass
self._status = "closed"
self._loading = False
if isinstance(e, ComputeError):
raise ControllerError(str(e))
else:
raise e
try:
os.remove(path + ".backup")
except OSError:
pass
self._loading = False
self.emit_controller_notification("project.opened", self.asdict())
# Should we start the nodes when project is open
if self._auto_start:
# Start all in the background without waiting for completion
# we ignore errors because we want to let the user open
# their project and fix it
asyncio.ensure_future(self.start_all())
async def wait_loaded(self):
"""
Wait until the project finish loading
"""
while self._loading:
await asyncio.sleep(0.5)
async def duplicate(self, name=None, reset_mac_addresses=True):
"""
Duplicate a project
Implemented on top of the export / import features. It will generate a gns3p and reimport it.
NEW: fast duplication is used if possible (when there are no remote computes).
If not, the project is exported and reimported as explained above.
:param name: Name of the new project. A new one will be generated in case of conflicts
:param reset_mac_addresses: Reset MAC addresses for the new project
"""
# If the project was not open we open it temporary
previous_status = self._status
if self._status == "closed":
await self.open()
self.dump()
assert self._status != "closed"
try:
proj = await self._fast_duplication(name, reset_mac_addresses)
if proj:
if previous_status == "closed":
await self.close()
return proj
else:
log.info("Fast duplication failed, fallback to normal duplication")
except Exception as e:
raise ControllerError(f"Cannot duplicate project: {str(e)}")
try:
begin = time.time()
# use the parent directory of the project we are duplicating as a
# temporary directory to avoid no space left issues when '/tmp'
# is located on another partition.
working_dir = os.path.abspath(os.path.join(self.path, os.pardir))
with tempfile.TemporaryDirectory(dir=working_dir) as tmpdir:
# Do not compress the exported project when duplicating
with aiozipstream.ZipFile(compression=zipfile.ZIP_STORED) as zstream:
await export_project(
zstream,
self,
tmpdir,
keep_compute_ids=True,
include_snapshots=True,
allow_all_nodes=True
)
# export the project to a temporary location
project_path = os.path.join(tmpdir, "project.gns3p")
log.info(f"Exporting project to '{project_path}'")
async with aiofiles.open(project_path, "wb") as f:
async for chunk in zstream:
await f.write(chunk)
new_project_id = str(uuid.uuid4())
# import the duplicated project
with open(project_path, "rb") as f:
project = await import_project(
self._controller,
new_project_id,
f,
name=name,
reset_mac_addresses=reset_mac_addresses,
keep_compute_ids=True
)
log.info(f"Project '{project.name}' duplicated in {time.time() - begin:.4f} seconds")
except (ValueError, OSError, UnicodeEncodeError) as e:
raise ControllerError(f"Cannot duplicate project: {str(e)}")
if previous_status == "closed":
await self.close()
return project
async def _fast_duplication(self, name=None, location=None, reset_mac_addresses=True):
"""
Fast duplication of a project.
Copy the project files directly rather than in an import-export fashion.
:param name: Name of the new project. A new one will be generated in case of conflicts
:param location: Parent directory of the new project
:param reset_mac_addresses: Reset MAC addresses for the duplicated project
"""
# remote replication is not supported with remote computes
for compute in self.computes:
if compute.id != "local":
log.warning("Fast duplication is not supported with remote compute: '{}'".format(compute.id))
return None
# work dir
p_work = pathlib.Path(location or self.path).parent.absolute()
t0 = time.time()
new_project_id = str(uuid.uuid4())
if location:
new_project_path = p_work.joinpath(location)
else:
new_project_path = p_work.joinpath(new_project_id)
# copy dir
await wait_run_in_executor(shutil.copytree, self.path, new_project_path.as_posix(), symlinks=True, ignore_dangling_symlinks=True)
log.info("Project content copied from '{}' to '{}' in {}s".format(self.path, new_project_path, time.time() - t0))
topology = json.loads(new_project_path.joinpath('{}.gns3'.format(self.name)).read_bytes())
project_name = name or topology["name"]
# If the project name is already used we generate a new one
project_name = self.controller.get_free_project_name(project_name)
topology["name"] = project_name
# To avoid unexpected behavior (project start without manual operations just after import)
topology["auto_start"] = False
topology["auto_open"] = False
topology["auto_close"] = False
# regenerate IDs for the duplicated project
regenerate_topology_ids(topology, new_project_path, reset_mac_addresses)
# dump the updated .gns3 project file
dot_gns3_path = new_project_path.joinpath('{}.gns3'.format(project_name))
topology["project_id"] = new_project_id
with open(dot_gns3_path, "w+") as f:
json.dump(topology, f, indent=4, sort_keys=True)
# update the snapshots with new IDs
snapshots_dir = os.path.join(new_project_path, "snapshots")
if os.path.isdir(snapshots_dir):
await update_snapshots(snapshots_dir, new_project_path, project_name, new_project_id)
os.remove(new_project_path.joinpath('{}.gns3'.format(self.name)))
project = await self.controller.load_project(dot_gns3_path, load=False)
log.info("Project '{}': fast duplicated in {:.4f} seconds".format(project.name, time.time() - t0))
return project
def is_running(self):
"""
If a node is started or paused return True
"""
for node in self._nodes.values():
# Some node type are always running we ignore them
if node.status != "stopped" and not node.is_always_running():
return True
return False
@open_required
def lock(self):
"""
Lock all drawings and nodes
"""
for drawing in self._drawings.values():
if not drawing.locked:
drawing.locked = True
self.emit_notification("drawing.updated", drawing.asdict())
for node in self.nodes.values():
if not node.locked:
node.locked = True
self.emit_notification("node.updated", node.asdict())
self.dump()
@open_required
def unlock(self):
"""
Unlock all drawings and nodes
"""
for drawing in self._drawings.values():
if drawing.locked:
drawing.locked = False
self.emit_notification("drawing.updated", drawing.asdict())
for node in self.nodes.values():
if node.locked:
node.locked = False
self.emit_notification("node.updated", node.asdict())
self.dump()
@property
@open_required
def locked(self):
"""
Check if all items in a project are locked and not
"""
for drawing in self._drawings.values():
if not drawing.locked:
return False
for node in self.nodes.values():
if not node.locked:
return False
return True
def dump(self):
"""
Dump topology to disk
"""
try:
topo = project_to_topology(self)
path = self._topology_file()
log.debug(f"Write topology file '{path}'")
with open(path + ".tmp", "w+", encoding="utf-8") as f:
json.dump(topo, f, indent=4, sort_keys=True)
shutil.move(path + ".tmp", path)
except OSError as e:
raise ControllerError(f"Could not write topology: {e}")
@open_required
async def start_all(self):
"""
Start all nodes
"""
pool = Pool(concurrency=3)
for node in self.nodes.values():
pool.append(node.start)
await pool.join()
@open_required
async def stop_all(self):
"""
Stop all nodes
"""
pool = Pool(concurrency=3)
for node in self.nodes.values():
pool.append(node.stop)
await pool.join()
@open_required
async def suspend_all(self):
"""
Suspend all nodes
"""
pool = Pool(concurrency=3)
for node in self.nodes.values():
pool.append(node.suspend)
await pool.join()
@open_required
async def reset_console_all(self):
"""
Reset console for all nodes
"""
pool = Pool(concurrency=3)
for node in self.nodes.values():
pool.append(node.reset_console)
await pool.join()
@open_required
async def duplicate_node(self, node, x, y, z):
"""
Duplicate a node
:param node: Node instance
:param x: X position
:param y: Y position
:param z: Z position
:returns: New node
"""
data = copy.deepcopy(node.asdict(topology_dump=True))
# Some properties like internal ID should not be duplicated
for unique_property in (
"node_id",
"name",
"mac_addr",
"mac_address",
"compute_id",
"application_id",
"dynamips_id",
):
data.pop(unique_property, None)
if "properties" in data:
data["properties"].pop(unique_property, None)
node_type = data.pop("node_type")
data["x"] = x
data["y"] = y
data["z"] = z
data["locked"] = False # duplicated node must not be locked
new_node_uuid = str(uuid.uuid4())
new_node = await self.add_node(
node.compute,
node.name,
new_node_uuid,
node_type=node_type,
**data
)
try:
await node.post("/duplicate", timeout=None, data={"destination_node_id": new_node_uuid})
except ControllerNotFoundError:
await self.delete_node(new_node_uuid)
raise ControllerError("This node type cannot be duplicated")
except ControllerError as e:
await self.delete_node(new_node_uuid)
raise e
return new_node
def stats(self):
return {
"nodes": len(self._nodes),
"links": len(self._links),
"drawings": len(self._drawings),
"snapshots": len(self._snapshots),
}
def asdict(self):
return {
"name": self._name,
"project_id": self._id,
"path": self._path,
"filename": self._filename,
"status": self._status,
"auto_start": self._auto_start,
"auto_close": self._auto_close,
"auto_open": self._auto_open,
"scene_height": self._scene_height,
"scene_width": self._scene_width,
"zoom": self._zoom,
"show_layers": self._show_layers,
"snap_to_grid": self._snap_to_grid,
"show_grid": self._show_grid,
"grid_size": self._grid_size,
"drawing_grid_size": self._drawing_grid_size,
"show_interface_labels": self._show_interface_labels,
"supplier": self._supplier,
"variables": self._variables,
"created_by": self._created_by,
}
def __repr__(self):
return f"<gns3server.controller.Project {self._name} {self._id}>"