Designing Agent-Assisted Abuse Detection With Classifier Routing and Feedback Loops

Designing Agent-Assisted Abuse Detection With Classifier Routing and Feedback Loops

An LLM agent will write you a convincing abuse report on almost any account you point it at. Give it tools to pull the account’s history and its connections, and a few minutes later you have a tidy paragraph explaining why this looks like a scam ring. The first time I watched one do it, I assumed the hard part was over. It wasn’t. Writing an explanation of why something looks abusive is usually the easy part. Whether an agent provides real value in production depends on the system around the agent: when it gets called, whether its findings ever improve conventional classifiers, whether anyone can explain a ban the agent recommended six months later, and what an attacker can do to circumvent agents. Abuse detection is the job of finding accounts, content, and coordinated campaigns that go against a platform’s rules and policies (fraud, scams, misinformation, etc.), all while malicious actors actively work to stay ahead of the abuse detection system. Start With Routing, Not Reasoning Picture a wave of freshly created accounts posting scam comments under popular videos. Any single account looks unremarkable - a genuine-looking name, a stock photo, a few posts. But the graph of accounts, the domains they link to, the devices they share, and the timing of their activity looks coordinated in a way no set of unrelated users ever would. That’s a good case for an agent: unfamiliar, relational, and worth a careful look. The mistake is making the agent the first stop for every suspicious account. Most suspicious accounts aren’t novel at all - they're the same abuse you've seen ten thousand times, and you already have classifiers that catch them. So before you build the agent, build the router that decides which cases the agent handles. Here’s roughly how I route them. (In what follows, “entity” just means the thing under investigation - an account, a piece of content, or a cluster of them.) classifier_score = known_abuse_model(entity) novelty_score = anomaly_model(entity) if classifier_score >= 0.95: send_to_known_abuse_queue() elif classifier_score = 0.90: send_to_agent_investigation() else: send_to_sampling_or_human_review() Treat those numbers as placeholders. I tune real ones against precision-recall curves and whatever reviewers can absorb, weighing what a missed scam costs against what a wrongful ban costs. The rule underneath it is simple: the agent runs only when the fast classifier is unsure and the entity is unusual. Everything else has a simpler solution. Known abuse belongs to a fast classifier - a gradient-boosted tree or a small transformer trained on behavioral features. This catches everything that resembles labeled abuse you've already seen, at a fraction of the latency and cost of an agent call. The novelty path asks a different question: how unusual is this entity compared to everything else you’ve seen? You can answer that with an isolation forest, with nearest-neighbor distance from known abuse clusters, or with community detection on a graph of accounts and the signals they share. The choice depends on your data: what matters is producing a single anomaly score per entity. The failure to watch for is the agent drifting into an expensive classifier. Running it on traffic a cheap classifier could have handled burns money for no gain, because an agent invocation costs far more than classifier inference. Make the Output Useful After the Investigation Here's a failure I've seen in practice. An agent investigates a cluster of malicious accounts and writes a clean, readable summary. A reviewer reads it and applies an enforcement action - a ban, a takedown, a rate limit. Two weeks later, a slightly modified version of the same attack shows up, and the entire process runs again from scratch. The agent did the work twice, but the classifier learned nothing either time. This is the most expensive failure mode I know of, and it almost always comes from a single design choice: the agent’s output is unstructured text instead of a structured training signal. A paragraph is something a human reads once and forgets. A structured record is something the rest of your system can actually consume. The agent should emit something closer to a training row: case_id: case_123 entities: account_1, account_2, domain_7 suspected_policy: scam classifier_score: 0.58 novelty_score: 0.96 distinguishing_features: - shared landing page domain - high outbound message fanout - recent account creation cluster agent_findings: - claim: accounts share infrastructure - evidence: graph_edge_1, domain_lookup_4 reviewer_label: confirmed_positive importance_weight: 0.8 Concretely, the agent emits the entity IDs involved, the features that separate those accounts from benign traffic, a label (confirmed positive, confirmed negative, or ambiguous) attributed to the process that assigned it, and the evidence trail behind that label. Attach a sample weight, and you have a row in the classifier's training set rather than a text report. The features record what made the case different; the evidence IDs make the label auditable long after anyone remembers the details.But you can’t retrain on all of it. You need a strategy for deciding which rows are worth learning from, because retraining on everything the agent flags may make the classifier worse, not better. Remember that the agent only ever sees cases that were already ambiguous, so the agent’s confirmed positives often cluster right around the decision boundary. Train on them indiscriminately and the model's precision collapses on the easy examples it used to get right. A better approach is to prioritize by how much a label would move the classifier. The most valuable rows are the ones your classifier got confidently wrong, and the agent will never find those on its own, because the router never sends them to it. They come from appeals and from sampling the auto-action queues. Convert those into the same record format so they land in the same training dataset. The endgame is distillation. Think of the agent as a slow, expensive teacher that explores, and the classifier as the fast student that has to operate at scale. Every confirmed novel pattern should move from teacher to student within a retraining cycle. If the agent is still catching a pattern six months after it first found it, the system’s distillation loop is broken. Make Every Agent Decision Reconstructable Sooner or later, someone asks why a specific account was banned. It might be a user appealing the decision, an internal review, a regulator, or your own team trying to debug a drop in precision. The question is always some version of: what made this case look like abuse? “The agent reasoned about it” is neither a defensible answer nor a useful one. LLM outputs aren't reproducible in practice, since the same prompt over the same evidence can produce different reasoning across model updates, infrastructure changes, or even retries inside a single deployment. Six months after a decision, you may not be able to reconstruct it at all. So treat every agent decision like a row in a system of record. At decision time, freeze the following signals: The evidence snapshot, exactly as the agent saw it. Log the full prompt (specific messages, accounts, and signals) as it existed in that moment. By the time anyone asks, the underlying data will have moved on. The model and prompt versions, including the version of every tool the agent could call. A tool whose backend changed last week is the same kind of state change as a prompt edit, and just as capable of making a decision unreproducible. The tool-call trace. This is what lets you reconstruct what the agent knew, as opposed to what it concluded. The abuse policy the case mapped to. Which specific rule did this fall under, and which features tied it there? You'll want this the moment the question becomes “is this consistent with how we handled similar cases?” In most production systems, the real attack surface isn't the model. It's the tools and the retrieved context around it. The moment your agent pulls in user-generated content as evidence, it's feeding adversarial input into its own context window. Anyone under investigation has every incentive to plant text designed to steer the investigation. You cannot eliminate that risk entirely, but you can reduce the blast radius. Start by scoping the abuse detection agent narrowly and capping what it can do per case. Budget the agent’s tool calls, retrieved rows and runtime per case, and block every tool the agent should not access. In the config below, the agent can read account age, message metadata, and the graph around an entity. It can't touch payment details or private profiles, and it can't call the enforcement tool at all. The agent can recommend a ban; it cannot execute one. So even if an attacker hijacks a single investigation via prompt injection, the damage is limited to one case rather than your whole user base. role: scam_investigator allowed_tools: account_age, message_metadata, domain_reputation, graph_neighbors blocked_tools: payment_details, private_profile, execute_enforcement limits: 12 tool calls, 30 seconds, 50 entities, no enforcement permission Retrieved user content should go into the prompt as quoted evidence. Most prompt-injection attempts are commands disguised as content: “ignore previous instructions and…”. Wrap retrieved content in a clearly demarcated part of the prompt that the model is told to treat as evidence, not direction. It’s imperfect - there's no true isolation between data and instructions inside an LLM's context window, but it raises the bar. The OWASP Top 10 for LLM Applications is a good read of the rest of this problem space. What this looks like in production Most traffic never reaches the agent. What the agent does find flows back into the classifier. Put it together and the structure of the system is unglamorous by design. A fast supervised classifier handles the large majority of decisions cheaply. An unsupervised novelty layer catches the unfamiliar cases the classifier isn't sure about. Agents investigate only the small fraction of cases that survives both filters, and emit structured records that improve the classifiers over time. Every step logs deterministically, tool access stays narrow and parameterized, and a human signs off on enforcement wherever agents or classifiers determine that the case is borderline. The interesting part - the agent and the reasoning it narrates so convincingly - is the smallest piece. It's also the piece that gets all the attention, which is why so many of these systems dazzle in a demo but disappoint months later.

Original Source

Read the full article at Hackernoon →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.