complylock.ai / aiuc-1-tool-call-testing

AIUC-1 tool-call testing

Tool-call testing asks one question the rest of your evaluation suite does not: when someone manipulates the conversation, what does the agent actually execute? ComplyLock runs your agent against a proprietary, continuously expanded adversarial corpus built for that surface, covering authority spoofing, privilege escalation through tool parameters, delayed exfiltration via planted instructions, and unsafe autonomous action chaining. Results are mapped to the relevant AIUC-1 controls where they apply, each with a baseline pass or fail, and every finding is reported by technique class and outcome only, because publishing payloads would retire the corpus. The output is the third-party evidence D004 expects, plus a ranked list of what to remediate first.

Book a call