Kindal
Blog / AI & Research
AI & Research

How to build an AI research agent that monitors sources

How to build an AI research agent that monitors sources
Key takeaways
  • A monitoring agent is a pipeline with a language model inside it, not a language model with tools attached. Most of the engineering is in the parts around the model.
  • Write the mandate first: what counts as relevant, what counts as material, and what the agent should ignore. Every later threshold is tuned against it.
  • Deduplicate on three levels: exact hashes for copies, near-duplicate fingerprints for rewrites, and event clustering for the same development told by different sources.
  • Persistent state is a set of explicit tables (items, events, reports, runs), not only a vector index. Retrieval finds similar text; it cannot say what was already reported.
  • Evaluate alert precision, misses found by audit, duplicate rate and quiet-day behavior. A run that sends nothing must still leave a log of what it read.
  • Put cheap filters before expensive ones: conditional requests, hashes and rules first, a small model for triage, a large model only for what survives.

The short answer

To build an AI research agent that monitors sources continuously:

  1. Write a mandate: what is relevant, what is material, what to ignore.
  2. Build a source layer from feeds, APIs and permitted pages, with health checks.
  3. Schedule each source by how often it publishes.
  4. Detect change and deduplicate items into events.
  5. Score each new event for relevance, novelty and materiality against the mandate.
  6. Summarize what clears the threshold, with a citation on every claim.
  7. Keep persistent state: items, events, reports and a log of every run.
  8. Deliver updates, and measure precision, misses and quiet-day behavior.

What is an AI research agent?

An AI research agent is software that uses a language model to plan research steps, gather sources, read them and write a cited result. Most agents in circulation run one investigation per request. A monitoring agent is the variant that keeps going after the first answer: it reads a fixed set of sources on a schedule and reports only what is new and relevant.

GPT Researcher is a well-known open-source example of the first kind. Its GitHub repository describes an Apache 2.0 licensed agent in which planner agents generate research questions, execution agents gather information, and a publisher aggregates the findings into a report with citations. It can search the web, read local documents and connect to other data through the Model Context Protocol. The README does not describe scheduling or change tracking, which is fair: that is not the job it was built for.

The difference between the two kinds is mostly outside the model:

  • A one-shot agent needs a planner, retrieval, a writer and citations.
  • A monitoring agent needs all of that, plus a scheduler, a memory of every item seen, deduplication, a relevance threshold, a record of what was already reported, and a way to end a run with nothing to say.

Treat a monitoring agent as a data pipeline with a language model at two or three points in it. Teams that start from the model and add plumbing later usually end up rebuilding the plumbing.

What are the components of a continuous monitoring agent?

A continuous monitoring agent has eight components, and each one hands a smaller, cleaner set of items to the next. The flow looks like this:

  1. Source layer. Fetches new items from feeds, APIs and pages.
  2. Scheduler. Decides which source to check, and when.
  3. Normalizer. Extracts the main text, title, author, date and canonical URL.
  4. Change detection and dedup. Drops copies, merges rewrites, groups items into events.
  5. Scorer. Rates each new event against the mandate.
  6. Writer. Turns events above the threshold into a cited update.
  7. State store. Holds sources, items, events, reports and run logs.
  8. Delivery and evaluation. Sends the update and records what happened.

In pseudocode, one run is short: for each source due now, fetch what is new; normalize it; discard what the store has already seen; attach the rest to existing events or open new ones; score the events that changed; if any clear the threshold, write and send an update; log the run either way.

The rest of this guide takes those steps one at a time.

How to build an AI monitoring agent, step by step

Step 1: Write the mandate before the code

Write the mandate as a short document before choosing any library. It is the specification every later threshold is tuned against, and it is the part most builds skip.

A useful mandate has four parts:

  • Scope. The subject, the entities that matter (companies, products, regulators, people) and their aliases.
  • Relevance. What kind of development the reader wants to know about. "Pricing changes, launches, executive departures and regulatory filings" is testable; "important news" is not.
  • Materiality. What makes a relevant item worth an interruption: a number crossing a level, a first occurrence, a change of direction.
  • Exclusions. What to ignore even when it matches: opinion pieces, event recaps, job ads, reposts.

Keep five to ten labeled examples next to the mandate, half that should alert and half that should not. They become your first test set in step 8.

Step 2: Build the source layer

Build the source layer from the most structured access a source offers, and fall back to page fetching only when nothing better exists. In order of preference:

  1. Official APIs, which give clean fields and stable identifiers.
  2. RSS and Atom feeds, which most blogs, newsrooms and newsletters still publish.
  3. Page fetching and extraction for changelogs, pricing pages and newsrooms without feeds.

For feeds, the Python library feedparser handles RSS and Atom and supports conditional requests. Its documentation on ETag and Last-Modified explains how to send those values back so an unchanged feed returns a 304 status with no body, and warns that skipping this causes repeated downloads that can get a client blocked. For pages, trafilatura is an Apache 2.0 Python package that extracts the main text and metadata such as title, author and date while dropping headers, footers and other boilerplate.

Respect the rules of each source. The Robots Exclusion Protocol is standardized as RFC 9309, which recommends that crawlers not rely on a cached robots.txt for more than 24 hours. The RFC 9309 text also states that these rules are not a form of access authorization, so a permissive robots.txt does not override a site's terms of service. Check both, identify your crawler with an honest user agent, and keep request rates low.

Give every source a health record: last successful fetch, last new item, error count. A source that fails silently looks exactly like a quiet topic.

Step 3: Schedule by source, not by topic

Schedule each source by its own publishing rhythm rather than running the whole topic on one clock. A regulator's docket and a startup's changelog do not move at the same speed, and polling both every ten minutes wastes requests on one and may still be too slow for the other.

A simple adaptive rule works well:

  • Start each source at a default interval, for example every few hours.
  • Shorten the interval when consecutive checks find new items.
  • Lengthen it, up to a ceiling, when checks keep coming back empty.
  • Pin a short interval for sources the mandate marks as time-critical.

Run fetches from a job queue with retries and backoff, and keep the scheduler separate from the processing workers. When processing falls behind, fetching should keep its pace, and the backlog should be visible rather than silently growing.

Step 4: Detect change and deduplicate

Detect change and deduplicate in three passes, from cheapest to most expensive, so that each pass sees fewer items than the last.

  1. Exact duplicates. Normalize the text (strip markup, whitespace, tracking parameters in URLs) and store a hash such as SHA-256. Identical hashes are the same item. This catches syndicated copies and feeds that re-publish old entries.
  2. Near duplicates. Lightly edited rewrites need a fingerprint that tolerates small differences. SimHash and MinHash are the standard choices. In the paper "Detecting Near-Duplicates for Web Crawling," Manku, Jain and Das Sarma of Google report that 64-bit SimHash fingerprints with a difference of at most 3 bits were reasonable for a repository of 8 billion web pages. The datasketch library implements MinHash with locality-sensitive hashing, which finds items above a Jaccard similarity threshold without comparing every pair.
  3. Events. Different sources describing the same development in different words are not near duplicates; they are one event. Cluster them by comparing embeddings of a one-sentence description of each item, then confirm the match with a structured check: do the key entities, numbers and dates agree?

The confirmation step is not optional. Embeddings place two different announcements by the same company close together because they share vocabulary, so a pure similarity threshold will merge distinct events. An entity and number check separates "the company raised prices" from "the company raised funding."

Each event gets a first-seen timestamp, a list of member items and a status: new, updated or already reported. Only new events, and material updates to reported ones, move on.

Step 5: Score relevance, novelty and materiality

Score each new or updated event on three separate dimensions, because they fail in different ways:

  • Relevance: does this event fall inside the mandate's scope and categories?
  • Novelty: does it add anything to what was already reported about this entity or thread?
  • Materiality: would the reader act differently because of it?

A practical implementation uses rules first and a model second. Rules handle the obvious cases cheaply: an excluded category, an entity not in scope, a source marked as low priority. A small language model then rates the remaining events against the mandate, with the labeled examples from step 1 in the prompt and a short written reason for each score.

Keep the three scores separate in storage rather than collapsing them into one number. When an alert turns out to be wrong, you can see which judgment failed. An event that was relevant and material but not novel points to a state problem; one that was novel but not material points to the threshold.

Add a daily cap on top of the threshold. A burst of coverage around one development should produce one update, not twelve.

Step 6: Summarize with citations

Write the update from the event records, not from raw search results, and attach a source identifier to every claim. The writer receives a compact package per event: the member items with their IDs, the prior report on the same thread if one exists, and the mandate.

Three rules keep summaries trustworthy:

  1. Every sentence cites at least one item ID, and IDs resolve to the original URL.
  2. A verification pass checks each citation against the cited text: the number, name or date in the sentence must appear in the source. Sentences that fail are rewritten or dropped.
  3. Attribution lives in the sentence. "The company's filing states" or "according to the regulator's notice" carries the status of a claim more precisely than a warning appended at the end.

Lead each update with what changed since the last report on that thread. The reader already has the background; the delta is the reason the message exists.

Step 7: Keep persistent state

Keep persistent state in explicit tables that the pipeline reads and writes on every run. A minimal schema:

  • Sources: URL, access method, schedule, health.
  • Items: canonical URL, hashes, fingerprint, extracted text, fetch time, source.
  • Events: description, entities, first-seen time, status, member items.
  • Reports: what was sent, when, to whom, which events and items it cited.
  • Runs: start and end time, items fetched, items dropped at each stage, events scored, whether anything was sent.

A vector index belongs alongside these tables, not in place of them. Postgres with the pgvector extension, which supports exact and approximate nearest neighbor search with HNSW and IVFFlat indexes, lets embeddings live next to the relational records in one database.

This is where retrieval alone falls short. A vector search returns text that resembles a new item. It cannot say whether the resembling event was reported, when, or to whom, and it cannot tell a restatement from an update. Those answers come from the events and reports tables. Retrieval is a lookup method; state is the record of decisions.

Step 8: Deliver, log and test

Deliver updates through the channel the reader already checks, such as email, Slack or a dashboard, and write a run log whether or not anything was sent.

The run log is what makes silence auditable. On a day with no update, it should show which sources were read, how many items arrived, how many were dropped at each stage and the highest score any event reached. With that record, "nothing happened" is a measured result rather than a guess.

Turn the labeled examples from step 1 into a regression test. Every change to a prompt, threshold or model runs against the set before it ships.

How do you evaluate a continuous monitoring agent?

Evaluate a monitoring agent on what it sends, what it misses and how it behaves when nothing happens. Four measures cover most of it:

  • Alert precision. Of the updates sent, how many would the reader have wanted? Review a sample every week. Falling precision is the earliest sign that the reader will stop opening the messages.
  • Misses found by audit. Recall is hard to measure because a miss is invisible. Approximate it by sampling: once a week, have a person scan a few sources directly and check whether anything material was skipped, and at which stage it was dropped.
  • Duplicate rate. How often does the same event reach the reader twice? A rising rate points to step 4 or to state that is not being read.
  • Quiet-day behavior. On days with no material change, the correct output is nothing. Count false alarms on quiet days separately, and check that the run log proves the sources were actually read.

Track latency too: the time from a source publishing to the reader receiving the update. For most mandates, an accurate update an hour later beats a noisy one in minutes.

How do you control the cost of a monitoring agent?

Control cost by putting cheap filters in front of expensive ones, so that the large model sees only the small fraction of items that survive. The order matters more than any single price:

  1. Conditional requests so unchanged feeds and pages cost one round trip and no download.
  2. Hashes and fingerprints to drop repeats before any model call.
  3. Rules for scope, exclusions and source priority.
  4. A small model for triage and scoring.
  5. A large model only for writing the final update.

Add a cap on items per run and per day, cache every summary by item hash so nothing is processed twice, and batch scoring calls where your provider allows it. Watch the run log for the ratio of items fetched to items reaching the writer. If that ratio rises, a filter upstream has stopped doing its job.

Be careful with features that add a model call per item, such as web search to enrich every event. They can quietly cost more than the rest of the pipeline combined.

What mistakes do builders make most often?

The mistakes that sink monitoring agents are almost all outside the model:

  • No written mandate. Without one, thresholds drift with whoever last tuned them, and precision cannot be measured.
  • Deduplicating by URL only. The same development arrives from ten domains; URL checks catch none of it.
  • Trusting embeddings to separate events. Similar vocabulary is not the same event. Confirm with entities and numbers.
  • Treating the vector index as memory. Without a reports table, the agent cannot know what it already told the reader.
  • Ignoring source health. A broken feed produces silence that looks like a quiet week.
  • Always sending something. A digest that arrives every day regardless of content trains the reader to skim it.

Should you build or buy a monitoring agent?

Build when monitoring is your product, when your sources are private or unusual, or when you need full control over the pipeline and its data. Buy when monitoring supports your work rather than being it, because the ongoing cost is maintenance: broken feeds, changed page layouts, model updates and threshold tuning.

If you build, an open-source research agent such as GPT Researcher can handle the deep-dive step, for example writing a background report when a new event opens a thread. The scheduler, state store, deduplication and evaluation described above still have to be written around it.

If you buy, judge products by the same architecture: do they read a source list you control, remember what they reported, stay silent on quiet days and cite every line? Kindal is one product built on this model: you choose the sources, it reads them and writes a brief only when something changes, with every line linked to its source. Teams comparing broader market intelligence options for small teams can apply the same test.

Whichever route you take, the measure is the same. A good monitoring agent is not the one that reads the most. It is the one whose messages you open, because each one tells you something you did not know and shows you where it came from.

Frequently asked questions

What is an AI research agent?

An AI research agent is software that uses a language model to plan and carry out research steps, such as searching, reading and summarizing sources, and returns a cited result. Most open-source research agents, including GPT Researcher, run one investigation per request: you ask a question, the agent searches and reads, and it writes a report. A monitoring agent is a variant that keeps running after the first report. It reads a fixed set of sources on a schedule, compares each new item with what it has already seen and reported, and sends an update only when something relevant and new appears. The second kind needs persistent state, deduplication and an explicit relevance threshold, which a one-shot agent can do without.

Is GPT Researcher good for continuous monitoring?

GPT Researcher is a strong starting point for the research step, but it is not a monitoring system on its own. According to its GitHub repository, it is an open-source, Apache 2.0 licensed agent that uses planner and execution agents to research a question across web and local sources and produce a cited report. Its README does not describe scheduling, change detection or a record of what earlier runs reported. To use it for monitoring, you would wrap it in a scheduler, store every item and report in a database, deduplicate new findings against past ones, and add a relevance threshold that lets a run end without output. Those parts are where most of the monitoring work lives.

How do you deduplicate news across many sources?

Deduplicate news in three passes, from cheapest to most expensive. First, normalize the text and compute an exact hash such as SHA-256 to drop identical copies and syndicated reposts. Second, use a near-duplicate fingerprint such as SimHash or MinHash to catch lightly edited rewrites of the same article. Third, cluster the remaining items into events: compare embeddings of a short description of each item, then confirm the match by checking that the key entities, numbers and dates agree. Embeddings alone tend to merge different events about the same company, so the entity check matters. Each cluster becomes one event with a first-seen time and a list of sources, and only new events can trigger an alert.

How do you stop an AI monitoring agent from sending too many alerts?

Stop an AI monitoring agent from over-alerting by giving it a written mandate, a scoring step and a hard threshold, then measuring precision. The mandate states what is relevant and what is material. The scoring step rates each new event on relevance to the mandate, novelty against past reports and materiality, and only events above the threshold reach the writer. Add a daily cap so one noisy day cannot flood the reader, and let the agent end a run with no output. Then review a sample of sent alerts each week and count how many the reader would have wanted. If precision drops, tighten the threshold or the mandate before adding sources.

How much does it cost to run an AI monitoring agent?

The cost of an AI monitoring agent depends mostly on how many items reach a language model, not on how many sources it reads. Fetching feeds and pages is cheap, especially with conditional requests that return no body when nothing changed. Hashing and rule-based filters cost almost nothing. The model calls are the main expense, so the design goal is to send as few items as possible to the largest model. A common pattern is a small, inexpensive model for triage, a larger model only for writing the final update, a cap on items per run, and caching so the same item is never summarized twice. With those controls, cost scales with the number of real developments rather than with source volume.

JH

Jonas Hale

AI & Research

Covers research agents, retrieval and the plumbing that makes machine reading useful. More interested in what fails than in what demos well.

Related articles