Incident Published: Sep 15, 2026

A Vision Model's Image Loader Became an Internal Port Scanner Twelve Hours After Disclosure

LMDeploy is a toolkit from Shanghai AI Laboratory for compressing, deploying, and serving language and vision-language models through an OpenAI-compatible API. Any http URL placed in a chat request's image_url field was fetched with no check on where it pointed, which is CVE-2026-33626. Sysdig saw the first exploitation attempt against its honeypot 12 hours and 31 minutes after the advisory published. In one eight-minute session the attacker used the image loader to probe AWS instance metadata, Redis, MySQL, and an admin interface.

Credential ExposureData ExfiltrationInput ValidationAI Infrastructure
7 applicable AIDEFEND defenses
Source: CVE-2026-33626: How attackers exploited LMDeploy LLM inference engines in 12 hours 
Author: Sysdig Threat Research Team
Original article: Apr 22, 2026

Threat Analysis

  • The vulnerable code was one unguarded fetch. When a chat message carried an image_url, LMDeploy's load_image() checked only that the string started with http, then retrieved it. With no hostname resolution, no private or link-local blocklist, and no redirect validation, the model server fetched any address the caller named.
  • The attacker used it as a probe, not an image request. Sysdig's honeypot recorded ten requests in three phases from 103.116.72[.]119. The first pointed the loader at the AWS instance metadata credentials path, then at Redis on loopback. The second fired a callback to a public out-of-band DNS service to confirm the request left the host, and read the OpenAPI schema. The third called the unauthenticated /distserve/p2p_drop_connect route, then swept localhost 8080, 3306, and 80 in 36 seconds.
  • Exposure came from defaults. The API server binds 0.0.0.0, and authentication is added only when an operator supplies API keys, so a default deployment answers anyone who can reach it.
  • Evidence boundary. The target was a honeypot. No source names a victim or states that credentials were retrieved; the impact language is conditional.

Applicable AIDEFEND Defenses (7)

AID-H-019.001
URL Normalization & Allowlist Filtering
Very High
Every server-side fetch triggered by caller-supplied input needs one safe wrapper that normalizes the URL, resolves the hostname, and refuses any address that is not publicly routable before a connection opens. That single gate is what the 0.12.3 fix added, and it is what turns the image loader back into an image loader instead of a general-purpose request primitive.
AID-H-003.010
Deployed AI Software Vulnerability Remediation Lifecycle
Very High
The fixed release was available roughly two weeks before the advisory, so the exposure window was a patching window rather than a zero-day. Reconcile running inference servers against ecosystem advisories by deployed image digest, set a deadline that assumes exploitation begins at disclosure, and rebuild to 0.12.3 or later rather than editing a live container.
AID-I-002.001
Internal AI Network Segmentation
High
Even a perfect request from the attacker only matters if the model server can reach something worth reaching. Placing inference workloads in their own segment with default-deny internal rules means that a fetch aimed at metadata services, Redis, MySQL, or an admin port has nowhere to land, which is the containment layer behind the fetch validation above.
AID-H-038.001
Inference Runtime Listener & Protocol Conformance
High
A serving process that binds 0.0.0.0 by default will quietly become internet-facing the first time it runs somewhere with a public interface. Record the approved bind addresses and ports for each runtime image, then read the actual listeners back after startup and fail the deployment when the process is listening more broadly than the profile allows.
AID-H-004.002
Service & API Authentication
High
LMDeploy adds its authentication middleware only when an operator passes API keys, so the default posture is an open inference API. Require an authenticated workload identity at the gateway or service boundary before any request reaches a model handler, and treat a serving endpoint with no configured caller identity as a deployment that should not be promoted.
AID-H-038.002
Inference Runtime Privileged Operation Authorization
High
The session's third phase called a distributed-serving control route that can tear down the link to a remote engine, and that route carried no authorization check of its own. Inventory every runtime route that changes execution or control-plane state, disable the ones a deployment does not use, and place an authenticated default-deny gate in front of those that remain.
AID-E-001.001
Root & Long-Lived Credential Object Eviction
Medium
No source confirms that credentials were retrieved, so this is a precaution rather than a response to proven theft. For any affected version that was internet-reachable, rotate the instance role credentials and any secrets the host could reach, because the requests observed were aimed squarely at the metadata path that issues them.

What Defenders Should Do Now

  • Inventory running inference servers by image digest, not by repository manifest, and upgrade LMDeploy to 0.12.3 or a later version. The same review should cover other serving stacks that fetch remote media on the model's behalf.
  • If a serving endpoint accepts remote image or document URLs, decide whether it needs to at all. Where it does, route every fetch through a wrapper that resolves the address and refuses loopback, link-local, and private ranges, and disable or revalidate redirects.
  • Check what your inference workloads bind to and who can route to them. A default 0.0.0.0 bind plus optional authentication left off is the combination that made this reachable in the first place.
  • Constrain the workload's own network position: deny it the cloud metadata endpoint, and require IMDSv2 with a hop limit of one so a server-side fetch cannot mint role credentials.
  • Add detection that does not depend on the application logging anything, such as alerting when an inference process connects to link-local, loopback, or internal service ports, and treat those alerts as high risk rather than noise.

Conclusion

Multimodal serving quietly hands a model server the ability to make outbound requests on a caller's behalf, and an image field is an unglamorous place for that capability to live. The attacker in this case did not need a novel technique; they needed a URL parameter, twelve hours, and a host that could still reach its own metadata service. AIDEFEND  points teams at the boundaries that decide the outcome: a safe fetch wrapper that refuses non-public addresses, a serving segment that leaves internal services unreachable, an exposure profile that catches a broad bind before it ships, and authenticated access to privileged runtime routes.