Research Published: Sep 13, 2026

SkillJack Turns Poisoned Agent Experience Into a Reusable Skill

The SkillJack paper evaluates SkillX and Anything2Skill, which transform agent experience trajectories into reusable skills. In a synthetic AppWorld study using one DeepSeek-v4-flash API, poisoned trajectories survived the transformation and evaded the evaluated routing-level detectors, with reported persistence after source records were deleted. This is a controlled research result, not a demonstrated compromise of a production skill registry or external service.

Model PoisoningSupply Chain DefenseAI Supply ChainAgentic AI
2 applicable AIDEFEND defenses

Threat Analysis

  • The durable artifact is the security boundary. The study does not stop at a poisoned prompt or one bad trajectory. It evaluates how experience becomes a reusable skill that can outlive the source records.
  • Transformation changes what detectors see. The paper reports a large difference between raw SkillX detection and the derived routing-level results. A detector that only inspects the input trajectory cannot be treated as an admission decision for the generated skill.
  • Persistence changes response requirements. The reported poisoned behavior remained after source records were deleted in the evaluation. Deleting the input is therefore not evidence that the derived artifact is safe.
  • The evidence is deliberately bounded. The evaluation uses a synthetic AppWorld setting, one DeepSeek-v4-flash API, and the paper's own detector and task conditions. It does not establish a compromised production registry, external account, or general result for every skill-generation system.

Applicable AIDEFEND Defenses (2)

AID-H-030.003
Manifest-vs-Observed Behavioral Consistency Testing
Very High
Before admitting a generated skill, compare a complete observed behavioral report with its declared permission manifest. This tests whether the trajectory-to-skill transformation added behavior that the artifact description does not disclose.
AID-H-030.005
Admission Decision Orchestration & Triggered Revalidation
Very High
Fuse behavior, provenance, and scanner evidence into a deterministic block, review, or allow decision bound to the exact skill artifact. Re-run admission when the artifact, policy, scanner, dependency intelligence, or evidence freshness changes.

What Defenders Should Do Now

  • Treat every generated skill as a new security artifact. Record its source trajectories, generator version, artifact digest, declared permissions, and intended execution scope.
  • Execute the candidate in an ephemeral restricted environment and compare the complete observed file, network, shell, tool, and resource behavior with the manifest before installation or reuse.
  • Make the final admission decision deterministic and fail closed. Missing, stale, malformed, insufficient, or mixed-artifact evidence must produce BLOCK or REVIEW, never an implicit allow.
  • Bind the decision to the exact artifact digest and re-run the full admission matrix when the artifact, publisher, scanner, policy, dependency findings, or evidence freshness changes. Keep the prior artifact unavailable while review is pending.

1 additional consideration

Research realism

SkillJack's reported rates are measured under a synthetic AppWorld setup with one model API and the paper's detector conditions. They should guide admission testing, not be presented as a fleet-wide attack probability.
Recommendation: Re-run behavior comparison and admission tests with the models, skill generators, registries, permissions, and runtime tools actually deployed by the organization.

Conclusion

SkillJack's important lesson is that a generated skill is a new trust object, not a transparent copy of its source trajectory. Defenders need behavior-versus-manifest testing and a digest-bound admission decision that is revalidated when evidence changes. AIDEFEND  provides those two controls without treating the paper's synthetic measurements as proof of production compromise.