Competitive intelligence sources: build a source database

- A source database records where competitive information comes from, not what you know: one row per source, with owner, type, URL or feed, what it reveals, cadence, reliability, legal notes, last check and signal history.
- The best competitive intelligence sources are specific to one competitor and change before any announcement: changelogs, docs, job boards, app store notes, filings, patents and partner listings.
- Every new source starts at F, reliability not yet judged, and earns a better grade from its own signal history rather than from its type.
- Track product changes with native feeds and APIs first, scoped page comparisons second, and full-page diffs only as a last resort.
- A quarterly review retires silent sources, re-rates the rest, repairs broken feeds and checks coverage competitor by competitor.
The short answer
A competitive intelligence source database is a registry with one row per source, saying what each source reveals about which competitor and whether it still earns its place. Build it in seven steps:
- Define the fields: owner, source type, URL or feed, what it reveals, cadence, collection method, reliability, legal notes, last checked, status, signal history.
- Run a discovery sweep for each competitor: domain, release trail, hiring, filings and patents, people and talks, ecosystem, customer voice.
- Rate each source, starting at F (not yet judged).
- Write the legal and ethical note before collecting anything.
- Pick a collection method per source: feed, API, scoped page comparison or manual check.
- Connect active rows to monitoring and log every useful signal.
- Review the registry every quarter: retire, re-rate, repair, check coverage.
What is a competitive intelligence source database?
A competitive intelligence source database is a structured list of every place your team collects competitor information from, with the metadata needed to judge each one. It answers "where does our information come from, and how good is it?" rather than "what do we know about this competitor?"
That distinction separates it from the documents it feeds. A competitor profile holds conclusions. A battlecard holds talking points. The source database sits underneath both, so that any claim in them can be traced to a row, and any row can be judged on its record.
Without one, sources live in bookmarks, alert settings and one analyst's memory, and the cost appears when that analyst leaves.
Why are the best competitive intelligence sources usually not news sites?
The best competitive intelligence sources are usually specific to one competitor and record change before anyone writes about it. News sites are neither: they cover many companies, and they report a change only once it is already a story.
For a registry, the practical test is specificity. A row for a competitor's changelog tells you exactly what you will learn and from whom. A row for a large business publication could produce anything about anyone, so it cannot be rated or pruned.
News still belongs in the registry, in a narrow role. Keep a few rows of type "discovery": a trade publication or a news query per competitor, read to find new sources rather than as a source of facts. When a story cites a filing, a job post or a partner announcement you were not tracking, that underlying record becomes a new row.
What fields should each source record have?
Each source record needs enough fields to answer four questions: what is this, what does it tell us, how do we collect it, and is it still worth it. Twelve fields cover that.
Identity
- Source ID. A short stable key, such as COMP-A-017, so signals and profiles can cite the row even after a URL changes.
- Competitor. The company the source covers, or "market" for sources that cover several.
- Source type. One value from a fixed list: website page, changelog, documentation, job board, filing, patent query, app listing, social account, video channel, partner listing, review site, community, discovery.
- URL or feed. The exact address you collect from. Where a machine-readable version exists (RSS, Atom, JSON API), record that instead of the human page.
Value
- What it reveals. One sentence on the kind of change this source shows first, for example "new integrations before they are announced" or "hiring for a second region." This field is the reason the row exists; if you cannot write it, the source does not belong.
- Update cadence. How often the source changes in practice: continuous, weekly, monthly, quarterly, irregular. Observed cadence beats the publisher's claim.
Collection
- Collection method. Native feed, API, page comparison (with the page section you compare), archive check or manual review.
- Owner. The person responsible for the row: they check it, rate it and decide when to retire it.
Quality and maintenance
- Reliability rating. A letter from A to F, explained in the reliability section below.
- Legal and ethical notes. Terms of service points, login requirements, personal data concerns, or "none."
- Last checked and status. The date of the last successful check, and active, watch or retired.
- Signal history. A running log of useful signals: date, one-line description, and whether it was confirmed or acted on.
To copy the template into a spreadsheet, paste this as the header row:
source_id, competitor, source_type, url_or_feed, what_it_reveals, update_cadence, collection_method, owner, reliability, legal_notes, last_checked, status, signal_history
Two hypothetical rows show the level of detail that works:
- COMP-A-004: Competitor A, job board, the public Greenhouse jobs endpoint for its board token, "new roles by team and location; first hire in a new function", continuous, API, product marketing lead, C, "public endpoint, no login; do not store applicant names", last checked this week, active, "three solutions engineer roles in Germany, confirmed by a partner launch six weeks later."
- COMP-A-011: Competitor A, app listing, its iOS App Store page, "release notes per version; features that ship to mobile first", weekly, page comparison on the release notes section, product manager, F, "none", last checked this week, watch, no signals yet.
How do you find the sources for each competitor?
You find a competitor's sources with a fixed discovery sweep, run the same way for every company, so that coverage gaps show up as empty categories rather than as surprises. Expect thirty to sixty minutes per competitor the first time.
Step 1: Start from the domain
Read the competitor's sitemap first. Most sites publish one at /sitemap.xml or list it in robots.txt, and the sitemap protocol allows an optional lastmod date per URL, so a sitemap is both a map of sections you did not know existed and a cheap way to spot new pages later.
Then check the usual subdomains and paths: docs, developers, status, trust or security, changelog, release notes, partners, customers, careers, investors. Look for RSS or Atom links in each section's page source, since blogs and changelogs often publish a feed without advertising it.
Step 2: Follow the release trail
The release trail is every place where shipped work is recorded:
- Changelog and release notes. The most direct product source; record the feed if one exists.
- Documentation and API reference. New endpoints, limits and deprecations often appear here before marketing names them.
- Code repositories. Public GitHub repositories publish releases, and each repository's releases page has an Atom feed at /releases.atom.
- App listings. The App Store and Google Play show notes for each version; Google's Play Console documentation describes a "What's new in this release?" field of up to 500 characters per language.
- Status page. Incident history reveals infrastructure, regions and components you would not otherwise see.
Step 3: Locate the hiring feed
Find which applicant tracking system hosts the competitor's careers page; the job links usually give it away. Several publish public, machine-readable boards. Greenhouse states in its Job Board API documentation that job board data is public and its GET endpoints need no authentication, with jobs listed under /v1/boards/{board_token}/jobs. Lever's postings API and Ashby's public job posting API offer similar read-only endpoints for published postings.
Record the API address, not the careers page. A careers page redesign breaks a page comparison; an API endpoint keeps working until the competitor changes systems, and that switch is itself worth noting.
Step 4: Add filings and patents
For public companies, SEC EDGAR offers an Atom feed of each company's filings, which can be filtered by form type, so current reports on material events arrive as a feed. Private companies that raise money under Regulation D usually file a Form D notice on EDGAR as well. For patents, save an assignee query in Google Patents or the USPTO's search tools; applications are slow but show where research money goes.
Step 5: Map people and talks
List the accounts of the founders, product leaders and developer advocates who speak for the company on X and LinkedIn, plus the company's video channel. YouTube channels have a public feed at /feeds/videos.xml?channel_id= followed by the channel ID, which turns uploaded webinars and conference talks into a collectable source. Add the agendas of the two or three conferences where the competitor reliably speaks.
Step 6: Check the ecosystem
Partners describe competitors in ways competitors do not. Check the marketplaces relevant to your category, such as AWS Marketplace, Salesforce AppExchange, the Shopify App Store or the Atlassian Marketplace, for listings, pricing notes and integration changes.
Step 7: Capture the customer voice
Add the competitor's pages on review sites such as G2, Capterra and Trustpilot, its public community forum, and the subreddits where its users talk. They rate lower on reliability but surface complaints before churn does.
Step 8: Add one discovery row
Finish with one news query or trade publication per competitor, typed as discovery. Its job is to point at records you missed in steps 1 to 7.
How do you rate the reliability of a competitive intelligence source?
Rate each source on its track record, with a scale that keeps the source's reliability separate from the credibility of any single item it publishes. The Admiralty code, used in NATO and Five Eyes intelligence doctrine and described in the US Army field manual FM 2-22.3, does exactly that.
The Admiralty code rates sources with letters:
- A: completely reliable
- B: usually reliable
- C: fairly reliable
- D: not usually reliable
- E: unreliable
- F: reliability cannot be judged
It rates each piece of information with a number, from 1 (confirmed by other sources) through 2 (probably true), 3 (possibly true), 4 (doubtful) and 5 (improbable) to 6 (truth cannot be judged).
For a CI registry, put the letter on the source row and the number on each entry in the signal history. A usually reliable source can still publish a doubtful item, such as a vague teaser post, and the split keeps that from contaminating the rating.
Three adaptations make the scale work for competitive sources:
- Every new source starts at F. Type gives you a prior, not a grade. A competitor's own filing will probably end up at A, but it earns the letter from its record.
- Rate accuracy and relevance together. A source whose signals are true but never matter, such as a social account that posts only event photos, should drift toward D for your purposes.
- Write down why. One line per rating change in the signal history ("moved to B: four of five signals confirmed this quarter") keeps the grade from becoming one person's impression.
What legal and ethical notes belong in the registry?
The legal and ethical notes field records any condition on how a source may be collected, written before the first collection. Most rows will say "none".
Typical notes cover:
- Terms of service. Some sites restrict automated access or reuse of their content. Note the restriction and choose the collection method accordingly.
- Logins and access controls. Anything behind a login you obtained honestly, such as a trial, needs a note on what the terms allow. Fake accounts, pretexting and posing as a customer do not belong in the registry at all.
- Robots.txt. The Robots Exclusion Protocol states that its rules "are not a form of access authorization." Treat robots.txt as the owner's stated preference for crawlers and follow it, but do not read it as permission: terms of service and law still apply.
- Personal data. Employee profiles, reviewer names and job applicants are personal data under many privacy laws. Collect the signal, not the person.
- Licensed sources. Paid databases often limit sharing; note who may see their output.
This field is guidance for the team, not legal advice. Rows with real ambiguity go to counsel before they go active.
How do you track competitor product changes automatically?
Track product changes automatically by choosing, for every release-trail row, the most structured collection method available: native feeds and APIs first, scoped page comparisons second, full-page comparisons last. Structure decides how much noise each method produces.
- Native feeds. Changelog RSS, GitHub releases Atom, EDGAR filings feeds and YouTube channel feeds deliver discrete, dated items.
- APIs. Job board endpoints return structured postings, so "new role" and "role removed" are exact comparisons rather than guesses.
- Sitemap comparison. Comparing the sitemap with last week's copy reveals new pages, such as a new solution page or a comparison page aimed at you, before they are linked anywhere.
- Scoped page comparison. For pricing pages, docs and app listings, compare only the section that matters, the plan table or the release notes, and compare text rather than HTML. Cookie banners, rotating testimonials and dates in footers are the main sources of false changes.
- Manual checks. Some sources resist automation; give them an owner and a calendar slot instead of a broken script.
When a method keeps producing false changes, record the fix in the row's collection field, not in someone's head.
How do you connect the registry to monitoring?
Connect the registry to monitoring by treating every active row as an instruction: collect from this address, with this method, at this cadence, and look for what the "what it reveals" field describes. The registry becomes the configuration of your monitoring, and the signal history its output.
Three links make it work:
- Rows to collection. Each active row maps to one job in whatever runs your monitoring, from a scheduled script to a dedicated platform. Teams comparing tools that monitor your market can use the registry as the requirements list: which source types a tool can collect, and by which method.
- Collection to triage. The "what it reveals" field tells whoever reads the output, a person or a model, what counts as a signal for that row. A changelog row that exists to catch integrations should not escalate a typo fix.
- Signals back to rows. Every signal someone confirms or acts on goes into the source's history. Without that loop, the quarterly review has nothing to measure.
Some teams delegate the reading layer entirely. Kindal reads the sources a team chooses, including company blogs, newsrooms, filings and patents, and writes a brief only when something in them changes, with every line linked to the source, which leaves the registry owner to judge the signal and log it.
How do you prune and review the registry every quarter?
Prune the registry every quarter with the signal history in front of you, row by row, so that decisions rest on what each source produced rather than on how important it sounds. A registry of a few hundred rows takes one to two hours.
Work through five checks:
- Retire the silent. A row with no useful signal for two quarters moves to retired, unless it is quiet by design. Filings feeds and patent queries can stay silent for months and still be essential; mark those rows so they are not cut by mistake.
- Re-rate. Move letters up or down from the quarter's record, with the reason logged.
- Repair. Fix broken URLs, redirected feeds and renamed board tokens. A competitor moving its careers page to a new applicant tracking system often signals a hiring push or a reorganization; log it as a signal before you update the address.
- Check coverage. For each important competitor, confirm at least one active source for product, pricing, hiring, partnerships and customers. An empty cell is a discovery task for next quarter.
- Rerun discovery for one competitor. Rotate through competitors so that every company gets a fresh sweep at least once a year.
Keep retired rows instead of deleting them: the history explains why the team stopped watching a source and stops someone adding it back.
What are the most common mistakes with a CI source database?
The most common mistakes come from treating the registry as a bookmark list rather than as a record that has to justify itself.
- No "what it reveals" field. Without it, every source looks equally useful and pruning becomes guesswork.
- Human pages instead of feeds. Tracking a careers page when an API exists multiplies false changes.
- No owner. Rows nobody owns go stale quietly, and nobody notices until a competitor's move is missed.
- Collecting before writing the legal note. Fixing a collection method after a terms of service complaint costs more than reading the terms first.
A source database is never finished. What matters is that every row can say what it is for, how good it has been and who looks after it.
Frequently asked questions
What are the main sources of competitive intelligence?
The main sources of competitive intelligence are the records competitors create themselves, plus the places where customers, partners and regulators describe them. Company-owned sources include the website, pricing page, changelog, documentation, blog, careers page and executive social accounts. Public records include securities filings, patent applications and regulatory submissions. Third-party sources include job boards, app store listings, partner marketplaces, review sites, community forums, conference agendas and trade press. Internal sources matter as well: sales call notes, win and loss interviews and support tickets often contain competitor details nobody publishes. The most useful sources are specific to one competitor and record change early, which is why changelogs and job boards usually beat general news sites for timing.
What should a competitive intelligence source list include?
A competitive intelligence source list should include, for every source, enough information to know why it is there and whether it still earns its place. The core fields are: a unique ID, the competitor it covers, the source type, the URL or feed address, what the source reveals, how often it updates, how it is collected, a reliability rating, legal or ethical notes, the date it was last checked, its status and a short signal history. The signal history is the field most lists lack. It records each useful change the source produced and when, so that a quarterly review can tell productive sources from noisy or silent ones with evidence instead of memory.
How do you rate the reliability of an intelligence source?
A common method is the Admiralty code, used in NATO and Five Eyes intelligence doctrine, which rates a source from A (completely reliable) to F (reliability cannot be judged) and, separately, rates each piece of information from 1 (confirmed by other sources) to 6 (truth cannot be judged). For competitive intelligence, keep the two ratings apart. The letter belongs to the source and lives in the registry. The number belongs to each signal and lives in the signal history. Start every new source at F and move it up or down based on its record: how often its signals were later confirmed, how early they arrived and how often they proved wrong or irrelevant.
Is it legal to collect competitive intelligence from public websites?
Collecting information that a competitor publishes openly, such as its pricing page, job posts, filings or release notes, is standard practice and generally lawful, but the method matters. Read each site's terms of service, respect access controls, and do not create fake accounts, impersonate customers or misrepresent who you are to obtain information. Robots.txt files signal what a site owner wants automated crawlers to do; following them is good practice even though they are not an access control. Be careful with personal data, such as employee profiles, which privacy laws may cover. Record these points in the registry's legal notes field, and ask counsel when a source sits in a gray area.
How often should you review competitive intelligence sources?
Review the source database as a whole every quarter, and check individual sources on the cadence at which they change. Fast sources, such as changelogs, pricing pages and job boards for your closest competitors, are collected continuously or weekly. The quarterly review is different: it looks at the registry itself. It retires sources that produced nothing useful for two quarters, unless they are quiet by design, re-rates reliability from the signal history, repairs broken URLs and moved feeds, and checks that every important competitor still has coverage for product, pricing, hiring, partnerships and customers. The review usually takes an hour or two for a registry of a few hundred rows.


