Your competitors already know about that RFP sitting on SAM.gov right now. By the time you see it, two or three other firms have spent months shaping the requirements, building relationships with the program office, and positioning their teaming arrangements. A knowledge graph built from federal procurement data (FPDS, SAM.gov, and USAspending) gives you a structural advantage by mapping agency buying patterns 6 to 12 months before solicitations appear. It turns procurement intelligence from a search problem into a relationship problem, and that distinction is worth millions in pipeline accuracy.
Most capture teams rely on GovWin alerts, Bloomberg Government searches, or keyword monitoring on SAM.gov. These tools tell you *what* is happening. They do not tell you *why* an agency buys the way it does, which contracting officers favor small business set-asides, which incumbents consistently win recompetes, or which prime-sub relationships signal a locked-up opportunity. That "why" lives in the connections between 4 million+ federal contract actions, and a knowledge graph is the only data structure designed to expose it.
The thesis is straightforward: if you connect three public federal data sources into a living, queryable graph and layer agentic ingestion pipelines on top, you will see recompete windows, teaming relationships, and agency procurement tendencies that your competitors miss entirely.
Your Competitors Already Know About the RFP. That Is the Problem.
Picture this. Your BD lead sends a Slack message at 8:47am: "New $45M IT modernization opportunity just posted on SAM.gov for DHS CISA. Responses due in 30 days." Your team scrambles. You pull the SOW, spin up a compliance matrix, and start identifying potential teaming partners.
Meanwhile, your top competitor identified this recompete 11 months ago by tracking the incumbent's period of performance expiration in FPDS. They attended the industry day in March. They submitted capability briefings to the program office in April. They locked in their small business sub by June. Your 30-day sprint is competing against their 11-month capture campaign.
Traditional competitive intelligence platforms are search tools. You type in an agency name or NAICS code and get a list of opportunities or awards. That is useful, but it treats every procurement event as an isolated record. The reality is that federal procurement is a network: agencies have contracting offices, contracting offices have officers, officers have award histories, awards have incumbents, incumbents have subcontractors, and subcontractors have past performance on adjacent contracts. When you model these relationships as a graph, patterns emerge that no keyword search can surface.
The structural gap in most capture operations is not access to data. FPDS, SAM.gov, and USAspending are all free and public. The gap is in how you structure and query that data. A relational database can tell you "Vendor X won 12 contracts from Agency Y." A knowledge graph can tell you "Vendor X has won 12 contracts from Agency Y, always with Vendor Z as a sub, always through Contracting Office W, and the next recompete window opens in 7 months."
Three Federal Data Sources That Build the Graph
Three federal data systems contain the raw material for a procurement knowledge graph. Each one contributes a different layer of intelligence.
FPDS-NG (Federal Procurement Data System - Next Generation) is the foundation. It records every contract action above the micro-purchase threshold: award amounts, NAICS codes, PSC codes, set-aside types, base and option periods, modifications, and the contracting office responsible. FPDS bulk data files are updated daily, and the system contains over 20 years of historical transactions.
SAM.gov serves double duty. Its entity registration side contains vendor capability narratives, socioeconomic certifications, and UEI numbers. Its opportunity side tracks the full solicitation lifecycle: pre-solicitation notices, sources sought, combined synopsis/solicitations, and award notices. Together, these tell you both who can compete and what they are competing for.
USAspending.gov adds the financial layer. It tracks obligation flows at the program activity and object class level, and (critically) it includes sub-award data. This is the only public source that reveals prime-sub relationships at scale.
| Data Source | Key Entities | Update Frequency | Access Method | Primary Capture Value |
|---|---|---|---|---|
| FPDS-NG | Contract actions, modifications, COs, vendors | Daily bulk files | Bulk download + API | Recompete timing, incumbent ID, set-aside patterns |
| SAM.gov | Entities, opportunities, exclusions | Near real-time | Public API (v2) | Solicitation lifecycle, vendor capabilities, UEI resolution |
| USAspending | Obligations, sub-awards, program activities | Monthly (sub-awards quarterly) | Bulk download + API | Prime-sub mapping, funding flow analysis, program-level spending |
The biggest technical challenge is entity resolution. The same vendor might appear as "Booz Allen Hamilton Inc." in FPDS, "BOOZ ALLEN HAMILTON INC" in SAM.gov, and "Booz Allen Hamilton" in USAspending. DUNS numbers are being phased out in favor of UEI, but historical records still use DUNS. Your entity resolution layer needs to handle name normalization, UEI/DUNS crosswalks, and parent-child corporate relationships. Get this wrong, and your graph will treat one vendor as three separate entities, which destroys every downstream query.
Knowledge Graph Architecture for Capture Teams (Not Data Scientists)
You do not need a PhD in graph theory to build this. You need a clear taxonomy of what matters to your capture pipeline.
Nodes are the things: agencies, contracting offices, contracting officers, vendors, contracts, NAICS codes, PSC codes, and program offices. Each node type has properties. A contract node carries award amount, period of performance dates, set-aside type, and competition status. A vendor node carries UEI, socioeconomic status, and NAICS capabilities.
Edges are the relationships: "awarded-to," "subcontracted-to," "competed-under," "modified-by," "succeeded-by," "set-aside-for." Each edge can carry properties too, like award date, modification number, or sub-award amount.
Why does a graph outperform a relational database here? Because the queries capture teams actually need are relationship queries. "Show me every vendor this contracting officer has awarded to in the last 5 years" requires multiple joins in SQL and gets progressively slower as data grows. In a graph database, that is a single-hop traversal that returns in milliseconds regardless of dataset size.
For technology choices, keep it pragmatic. Neo4j Community Edition is free and has the largest ecosystem, but its AGPL license may be a concern for proprietary deployments. Amazon Neptune is a managed service that eliminates infrastructure overhead at roughly $0.10 per hour for a db.r5.large instance, making it viable for mid-size capture shops. Apache TinkerPop with JanusGraph gives you a fully open-source stack but requires more engineering to stand up. If your team has a data engineer, start with Neo4j. If you have an AWS account but no graph experience, Neptune with Gremlin queries will get you running in days.
The Agentic Layer: Continuous Ingestion, Not Monthly Reports
The word "agentic" gets thrown around loosely. Here it means something specific: autonomous software agents that ingest, reconcile, enrich, and alert on procurement data without requiring a human to trigger each step.
Key Statistics
4.2M+
Federal contract actions in FPDS over the past 5 fiscal years, each one a potential signal
6-12 mo
Lead time advantage when tracking recompete signals vs. waiting for SAM.gov postings
23 min
Median time from FPDS delta ingestion to alert generation in a properly configured pipeline
87%
Of recompete opportunities that show at least one detectable signal in FPDS modification history before RFP release
A static dashboard that refreshes monthly (or worse, quarterly) misses critical signals. Contract modifications happen daily. A bridge contract award or a short period-of-performance extension is a flashing neon sign that a recompete is imminent, but only if you see it within days, not weeks.
Your agentic layer should include four agent types:
- Entity Resolver Agent: Matches new FPDS and USAspending records to existing graph nodes, creates new nodes when no match exists, and flags ambiguous matches for human review
- Contract Lifecycle Tracker: Updates contract nodes with modifications, calculates remaining period of performance, and tags contracts approaching option period decisions
- Recompete Countdown Agent: Monitors contracts within 18 months of PoP expiration, tracks bridge contract indicators, and fires alerts when recompete probability crosses a configurable threshold
- Anomaly Detector: Flags unusual patterns like a contracting office switching from sole-source to full-and-open competition, a sudden increase in small business set-asides for a product service code, or an incumbent receiving an unusually short extension
Here is a practical example. Your anomaly detector agent notices that contract FA8732-21-C-0045, a $12M IDIQ for cybersecurity services at an Air Force installation, just received a modification extending the period of performance by only 6 months. The original contract had 12-month option periods. A 6-month extension is a bridge, which means the contracting office is already planning a recompete but needs time to finalize the solicitation. Your pipeline flags this, enriches it with the incumbent vendor's sub-award history and the contracting officer's award patterns, and pushes an alert to your capture manager. That alert arrives 6 to 9 months before the RFP hits SAM.gov.
Mapping Incumbent and Teaming Relationships Your BD Team Cannot See
Sub-award data from USAspending is the most underused public dataset in federal capture management. It reveals prime-sub partnerships that never appear in press releases, teaming announcements, or LinkedIn posts.
The Single Most Valuable Graph Query for Capture Teams
Query: "For target agency X and NAICS code Y, show me all prime-sub pairs that have appeared together on 2+ contracts in the last 5 years, ranked by total combined award value." This one query exposes entrenched teaming relationships, identifies which small businesses are already locked up as subs, and reveals which primes might be looking for new partners because their usual sub just won a prime contract elsewhere.
The pattern is reliable. When Vendor A primes and Vendor B subs on three or more contracts for the same agency within the same NAICS code, they are almost certainly teaming on the next recompete. If your BD team plans to approach Vendor B as a teaming partner without knowing this relationship exists, they will waste a quarter of pursuit effort before getting a polite "we already have a prime."
Negative signals are equally valuable. When an incumbent loses a recompete, that is data. When a contracting office shifts from sole-source to full-and-open, that is data. When a historically loyal agency-vendor relationship shows no new awards in 18 months, something has changed. Your graph should surface these breaks in pattern as loudly as it surfaces the patterns themselves.
Consider this scenario. Your team is targeting an OASIS+ task order for cloud migration services at the Department of Veterans Affairs. A graph query reveals that the likely incumbent, a mid-tier firm, has used the same small business HUBZone sub on four consecutive VA contracts. But that HUBZone sub just won their own prime contract with the Army, which will consume most of their bench capacity. The incumbent now has a teaming gap. Your firm approaches the incumbent as a replacement sub, or you approach the HUBZone sub's competitors who are now viable alternatives. Either way, you identified this opening 9 months before the task order dropped because the relationship data was in the graph.
Recompete Timelines: The Highest-Value Signal in the Graph
Recompete timing is the single highest-value output of a procurement knowledge graph, and most capture teams calculate it wrong.
The common mistake: tracking the original award date and adding the base plus option periods to estimate when the contract ends. The problem is that modifications change everything. Extensions, scope additions, bridge contracts, and option exercises all shift the effective end date. If you are working from the original award in GovWin and have not checked the latest FPDS modification, your recompete estimate could be off by 12 to 24 months.
The correct approach: pull every modification for a contract from FPDS, identify the latest period of performance end date, check for bridge indicators (short extensions, J&A postings, sources sought notices), and calculate a confidence-weighted recompete window.
| Signal Type | Data Source | Confidence Level | Recommended Capture Action |
|---|---|---|---|
| PoP expiry within 18 months, no extension filed | FPDS modification history | High (80%+) | Begin shaping, schedule agency meetings |
| Bridge contract (extension < 50% of normal option period) | FPDS modification + anomaly detection | Very High (90%+) | Accelerate capture, finalize teaming, draft capability brief |
| Sources sought or RFI posted | SAM.gov opportunity feed | Near Certain (95%+) | Full capture sprint, assign proposal manager |
| J&A (Justification & Approval) posted for current contract | SAM.gov + FPDS sole-source flag | High (85%+) | Monitor for full-and-open shift, prepare competitive response |
| Funding reduction in program activity | USAspending obligation data | Moderate (60%) | Assess whether scope will shrink, adjust bid ceiling accordingly |
Build a rolling 18-month recompete calendar filtered by your target NAICS codes, agencies, and minimum contract ceiling. Update it daily from your agentic pipeline. Connect each recompete entry to your pipeline stages so BD leadership sees forecast revenue impact, not just a list of contracts. When a recompete moves from "estimated" to "bridge detected," that should automatically escalate the opportunity's priority in your capture pipeline and pursuit gate process.
From Graph Query to Capture Decision: Making This Actionable
A knowledge graph that answers questions is useful. A knowledge graph that drives go/no-go decisions is valuable. The gap between the two is translation.
Your graph can calculate an agency loyalty index: how often does this contracting office recompete to the same incumbent versus awarding to a new vendor? If the loyalty index is above 80%, you are fighting uphill as an outsider unless you have a specific discriminator or the incumbent has performance issues.
It can calculate competition density: how many unique vendors have won contracts in this NAICS/agency intersection in the last 3 years? A density of 2 to 3 means an entrenched competitive set. A density of 8+ means the agency is willing to spread awards.
It can surface set-aside pattern shifts: has this contracting office increased its proportion of small business set-asides over the last 24 months? If your firm is a large business, that is a signal to pursue as a sub or mentor-protege prime, not a prime.
These outputs map directly to your existing go/no-go framework:
| Graph Output | Capture Decision Input | What It Tells You |
|---|---|---|
| Agency loyalty index: 85% | Incumbent advantage score | This CO rarely switches vendors; you need a strong discriminator or a teaming arrangement with the incumbent |
| Competition density: 3 vendors in 5 years | Competitive landscape assessment | Tight competitive set; identify what differentiated the winners |
| Set-aside trend: 40% to 65% SB over 24 months | Bid vehicle selection | Pursue through your small business partner or JV, not as a large prime |
| Recompete window: 7 months, bridge detected | Pipeline stage and urgency | Escalate to active capture, assign proposal manager now |
| Sub-award pattern: incumbent's usual sub is unavailable | Teaming opportunity score | Approach the incumbent as a replacement sub, or compete directly with your own SB partner |
When these graph-derived scores feed into your compliance matrix and proposal structure, your team starts every pursuit with a data-backed assessment instead of gut instinct and stale market research.
Build This in 90 Days, Not 18 Months
You do not need a year-long data engineering project. You need a focused 90-day build with clear phase gates.
Phase 1 (Weeks 1 to 4): Foundation. Download FPDS bulk data for the last 5 fiscal years, filtered to your top 3 target agencies. Normalize vendor names and resolve UEI/DUNS crosswalks. Stand up a Neo4j instance (Community Edition is fine for v1) with three node types: Agency, Vendor, Contract. Load the data and run your first queries. Deliverable: a working graph you can query for "show me all vendors awarded contracts by [target agency] in NAICS [your code]."
Phase 2 (Weeks 5 to 8): Relationships and Signals. Integrate USAspending sub-award data to build prime-sub relationship edges. Implement recompete countdown logic using PoP end dates and modification history. Build your first alert pipeline: a daily script that checks for new FPDS modifications on tracked contracts and sends alerts via email or Slack. Deliverable: a recompete calendar with confidence scores and a functioning daily alert system.
Phase 3 (Weeks 9 to 12): Enrichment and Integration. Add SAM.gov entity data for capability narratives and socioeconomic certifications. Implement relationship scoring (agency loyalty index, competition density, sub-award patterns). Connect the graph outputs to your capture management workflow, whether that is a spreadsheet, a CRM, or a purpose-built tool like Projectory. Deliverable: a weekly capture briefing powered by live graph intelligence.
What to skip in v1: do not try to build NLP-based requirement extraction from SOWs. Do not attempt to parse unstructured documents. Start with structured data only. The structured data in FPDS, SAM.gov, and USAspending contains enough signal to transform your capture process. You can add NLP layers in v2 after you have proven the value of the graph.
Frequently Asked Questions
How much does it cost to build a procurement knowledge graph?
Your primary costs are engineering time, not infrastructure. Neo4j Community Edition is free. AWS Neptune runs about $0.10/hour for a small instance. The real investment is 60 to 120 hours of data engineering over 90 days to build ingestion pipelines and entity resolution. For a mid-size GovCon firm, that is roughly $15K to $30K in loaded labor cost, far less than a single GovWin enterprise license.
Is FPDS data really updated daily?
Yes. FPDS publishes daily delta files containing all contract actions processed in the previous 24 hours. The bulk historical archive is updated monthly. For a live capture intelligence system, you want both: the historical archive as your baseline and daily deltas to keep the graph current.
Can a small capture team (3 to 5 people) actually use a knowledge graph?
Absolutely. The graph replaces manual research, it does not add to it. Instead of spending 4 hours per opportunity researching incumbents and agency history, your team queries the graph and gets answers in minutes. The 90-day build requires a data engineer or a technically capable analyst, but ongoing use is point-and-click query execution.
What about classified or controlled procurement data?
This approach uses only publicly available, unclassified data. FPDS, SAM.gov, and USAspending are all public by law. Some contract actions related to classified programs are redacted or excluded from FPDS, so your graph will have gaps for intelligence community and certain DoD programs.
Summary: Your Concrete Next Steps
The structural advantage in federal capture management is not more data. It is connected data. A knowledge graph built from FPDS, SAM.gov, and USAspending turns isolated procurement records into a queryable map of how agencies buy, who they buy from, and when the next buying decision is coming.
Here is what to do this week:
- Download FPDS bulk data for your top 3 target agencies, covering the last 5 fiscal years. The files are free at fpds.gov.
- Count the recompetes you missed. Filter for contracts with PoP end dates in the last 12 months. How many of those did your team identify before the solicitation posted? If the answer is less than half, your current CI process has a structural gap.
- Start tracking one metric: time from first detectable recompete signal (PoP expiry, bridge contract, sources sought) to your team's awareness of the opportunity. If that number is greater than 30 days, you are ceding months of capture positioning to competitors who are watching the data more carefully.
Remember the scenario from the opening. Your competitor saw the DHS CISA opportunity 11 months before you did. That head start did not come from better relationships or a bigger BD team. It came from watching the data, connecting the relationships, and acting on the signals while your team was still waiting for SAM.gov to tell them what was already in motion.