Analyzing model input and output logs in an AI-native detection pipeline to understand and uncover malicious AI agent behavior