It is Thursday afternoon, the volume is due Monday, and your proposal manager just asked the same question she asks on every pursuit: "Do we have a good past performance for cloud migration under a firm fixed price vehicle?" Somebody remembers a 2023 project. Somebody else thinks the CPARS rating was Very Good but is not sure. The person who actually ran it left the company in March. Now three people spend the weekend reconstructing a story that existed in complete form eighteen months ago.
The fastest way to fix this is to stop reconstructing past performance under deadline and start assembling it continuously. An agentic AI system does this by ingesting deliverables, CPARS ratings, and client feedback into a scored evidence store, then auto-drafting narratives that are already ranked against an RFP's relevancy and recency criteria before a human opens the volume. Done right, this cuts volume production time by roughly 60% and shifts human effort from writing to validating claims and CPARS traceability.
This article walks through the four cooperating agents that make it work, how to build and tag the evidence store, how to score drafts against the evaluation rubric, and how to keep everything defensible. I will also give you a 30 day rollout and a first action you can take in the next half hour.
The Weekend Nobody Should Spend Rebuilding Past Performance
Every proposal shop I have worked in treats past performance as a search-and-rescue mission. The solicitation drops, the capture lead identifies three to five references, and a writer starts hunting for the raw material: the final deliverables, the CPARS narrative, the customer point of contact, the dollar value, the period of performance. That material lives in six different places, and half of it is only in someone's head.
The real cost is not the hours. It is the institutional knowledge that evaporates. When the program manager who ran a $40M IT modernization contract leaves, the nuance of that engagement leaves with her. The next writer reconstructs an approximation, misses the specific outcomes the customer valued, and produces a narrative that scores lower than the work actually deserved.
Evidence should be captured while the project is running and the memory is fresh, not reassembled under a Monday deadline. That is the shift agentic orchestration makes possible. Instead of a person hunting records for each bid, a set of agents continuously pulls, tags, and scores your performance history so it is bid-ready the moment a solicitation arrives.
Why Manual Past Performance Volumes Fail Evaluators
Evaluators do not score prose quality. They score against the factors in Section M, and for past performance those factors are almost always relevancy, recency, and demonstrated quality of performance as documented in sources like CPARS. A beautifully written narrative about an irrelevant contract earns nothing.
Manual narratives fail in three predictable ways. First, they miss the specific alignment the RFP demands. The solicitation asks for experience in a particular PWS task area, and the writer submits a reference that is adjacent but not on point. Second, the claim language goes stale. A metric that was accurate in the 2022 proposal gets copied forward without verification, and now it does not match the CPARS record. Third, unverifiable numbers create evaluation risk. When an evaluator cannot confirm a claim against a source, the confidence in the entire submission drops.
The Contractor Performance Assessment Reporting System exists precisely so evaluators can verify quality of performance across the government [3]. If your narrative claims outcomes that the CPARS record does not support, you are inviting the evaluator to discount you.
Key Statistics
28%
Share of a proposal professional's time spent searching for and reconciling reusable content rather than writing [1]
41%
Portion of federal contractor CPARS evaluations that carry no narrative detail beyond the rating, forcing teams to reconstruct context manually [2]
19,000+
New CPARS assessments entered into the system in a single fiscal year, the volume of records teams must track [3]
The lesson: manual assembly optimizes for the wrong thing. Writers polish sentences while the evaluator is checking whether the reference is relevant, recent, and verifiable. An agentic pipeline inverts this by getting the evidence and scoring right first, then treating prose as the last mile.
The Four Agents That Do the Work
An agentic system is not one large model doing everything. It is a set of specialized agents, each with a narrow job, that hand work to each other. For past performance, four agents cover the workflow.
The ingestion agent connects to your source systems and pulls the raw material: final deliverables, CPARS narratives and ratings, client feedback emails and surveys, invoices, and contract modifications. It normalizes these into a single evidence store and records where each item came from and when it was last verified.
The mapping agent tags every evidence item with the metadata evaluators care about: NAICS code, contract type (FFP, T&M, CPFF), total contract value, period of performance, prime or subcontractor role, and capability areas. This tagging is what lets the system answer "cloud migration, FFP, within the last three years" in seconds instead of a weekend.
The drafting agent takes a target RFP's evaluation factors and generates narrative sections aligned to them. It pulls only from tagged, verified evidence and quotes stored metrics rather than inventing them. This is the same discipline that makes a good content reuse library work: the source material is structured, so the draft is traceable.
The evaluator-simulation agent scores each draft against a reconstructed relevancy and recency rubric before any human reads it. It flags weak references, missing CPARS linkage, and stale claims, then routes only high-confidence drafts to writers. The rest go back to capture with a specific gap to fill.
These agents cooperate. Ingestion feeds mapping, mapping feeds drafting, and the evaluator agent grades the output and closes the loop by flagging what is missing.
Building the Evidence Store: What Goes In and How It Is Tagged
The evidence store is the foundation. If the records are incomplete or untagged, every downstream agent produces weak output. Define a canonical evidence record with required fields so nothing enters the store half-formed.
Each record should carry: a unique contract identifier, the source system it came from, a freshness date (when the underlying fact was last verified), a verification status (verified, unverified, or disputed), and the CPARS entry it links to. That last field matters most. A claim that cannot point to a CPARS rating or a signed customer document is a claim you should not put in front of an evaluator.
| Evidence Type | Key Tags Applied | CPARS Linkage | Freshness Rule |
|---|---|---|---|
| Final deliverable | Capability area, PWS task, contract value | Links to rating period | Flag if PoP ended over 3 years ago |
| CPARS narrative + rating | Rating tier, evaluation factor, agency | Primary source of record | Flag if superseded by newer rating |
| Client feedback / survey | Sentiment, named outcome, POC | Corroborates rating | Flag if older than 12 months |
| Invoice / mod | Dollar value, contract type, PoP | Confirms scale and duration | Flag if not matched to a deliverable |
Handling Controlled Unclassified Information (CUI) is non-negotiable in this pipeline. Deliverables and client feedback frequently contain CUI, so the ingestion agent must enforce access controls and encryption consistent with NIST SP 800-171 requirements [4]. In plain terms: the store must restrict who can see each record, log access, and keep protected data segregated so a drafting agent never surfaces controlled content into a document that goes to the wrong audience.
Freshness scoring is what keeps stale claims out. Give every record a freshness score based on its verification date and the end of the period of performance. When a record ages past the RFP's recency window, typically three years for relevancy, the system flags it automatically instead of letting a writer copy it forward by habit.
Scoring Drafts Against the Evaluation Rubric Before Humans See Them
The evaluator-simulation agent is the piece that separates this approach from ordinary automation. It reconstructs the RFP's relevancy and recency criteria into a machine-readable rubric, then grades each draft the way a source selection evaluator would.
Start by parsing Section L and Section M. Relevancy usually breaks into scope, magnitude of effort, and complexity. Recency is a date window. The agent converts these into scored checks: does this reference match the required scope, is the dollar value comparable, did the work occur within the recency window, and is there a CPARS rating to support the quality claim. This is the same logic behind compliance matrix automation, applied to evidence scoring rather than shall-statement coverage.
Catch Weak Evidence Before It Costs You a Point
Run evaluator simulation the day the solicitation drops, not the week the volume is due. If the agent flags that your strongest cloud reference ended 40 months ago and falls outside a 36 month recency window, capture has time to substitute a fresher reference or request a customer letter. Discovering that gap during pink team review is too late to fix.
Here is what a scored narrative element looks like in practice:
Claim: Migrated 1,200 workloads to FedRAMP-authorized cloud
CPARS ref: Contract N00178-21-C-3042, Rating: Very Good (Quality)
Relevancy: 0.91 (scope match, magnitude within range)
Recency flag: PASS (PoP ended 14 months ago, inside 36-month window)
Verification: VERIFIED (deliverable + CPARS + customer survey)Only drafts that clear a confidence threshold, say relevancy above 0.80 with a passing recency flag and verified status, route to writers. Everything else goes back to capture as a specific, named gap. That routing decision is the difference between a writer polishing a strong draft and a writer discovering on Friday that the reference does not hold up.
The 60% Time Cut: Where the Hours Actually Disappear
The time savings are not magic. They come from eliminating the search-and-reconstruct phase and shrinking the writing phase to validation.
| Volume Step | Manual Workflow | Agentic Workflow | Time Impact |
|---|---|---|---|
| Find references | Days of hunting across systems and people | Pre-tagged store returns matches in minutes | Near elimination |
| Verify CPARS + metrics | Manual lookup, often skipped under pressure | Auto-linked at ingestion, flagged if stale | Hours to minutes |
| Draft narratives | Writer builds each from scratch | Agent drafts from verified evidence | Roughly 50% cut |
| Score against Section M | Done late, during review, if at all | Scored before human sees draft | Shifts left, prevents rework |
| Human review | Writing plus fact-checking under deadline | Validating traceability and claim language only | Focused, faster |
The human effort shifts from producing prose to confirming that every claim traces to a verified source and approving the claim language. That is a fundamentally faster task. Validating a well-sourced draft takes a fraction of the time it takes to build one from nothing.
Reuse compounds the savings. The first pursuit populates and tags the evidence store. Every subsequent pursuit starts from that pre-scored base, so the marginal cost of each new past performance volume drops with each bid. A shop that runs a dozen pursuits a year sees the store get richer and the scoring get sharper over time.
Consider a multi-volume pursuit requiring five past performance references across two capability areas. Manually, that is easily 60 to 80 hours split across writers and capture, much of it in reconstruction. With a populated evidence store and evaluator simulation, the same work compresses to roughly 25 to 30 hours, concentrated in validation and refinement. That is where the 60% figure comes from: the reconstruction and blind-drafting hours simply stop existing.
Guardrails: Keeping Agentic Narratives Defensible and Truthful
Speed means nothing if the narrative is not defensible. The whole system has to hold up to a customer's scrutiny and your own internal review, so build the guardrails in from the start.
Every generated claim must trace to a verifiable source record and a CPARS entry. If the drafting agent cannot cite a stored, verified fact, it does not write the sentence. This is the single most important rule, because it eliminates the failure mode that worries every compliance analyst: a confident-sounding metric that no one can back up.
A human sign-off gate is mandatory before any narrative enters the final volume. The agents assemble, score, and draft, but a person approves. This is not a formality. It is where a subject matter expert confirms the story matches reality and the claim language is one the company will stand behind.
Maintain an audit trail. For each narrative, keep the chain from claim to source: which deliverable, which CPARS rating, which customer document. If the customer challenges a claim during evaluation, or if your own past performance review board wants to confirm accuracy, the trail answers the question in seconds. Federal guidance on AI use emphasizes traceability and human accountability for exactly this reason [5].
Finally, forbid hallucinated metrics by design. The drafting agent quotes only numbers that exist in the store with a verification status of verified. It never estimates, rounds up, or infers a metric to make a sentence stronger. A slightly less impressive claim that is true always beats an impressive claim that collapses under verification.
Your First 30 Days: Standing Up the Pipeline
You do not need to build all four agents in week one. Start with the evidence store, because everything else depends on it.
- Week 1 - Inventory: Pull your last two to three years of contracts and list, for each, the final deliverable location, CPARS rating and narrative, customer POC, contract type, value, and period of performance. This surfaces your gaps immediately.
- Week 2 - Build the tagging schema: Define the canonical evidence record and tags (NAICS, contract type, capability area, PoP, verification status). Apply CUI access controls consistent with NIST SP 800-171 before loading anything sensitive.
- Week 3 - Pilot one capability area: Populate the store for a single capability where you bid often. Add freshness scores and CPARS linkage. This keeps the pilot small enough to finish and real enough to prove value.
- Week 4 - Add scoring: Reconstruct one recent RFP's relevancy and recency criteria into a rubric and run the evaluator simulation against your pilot evidence. Compare its scores to how those references actually fared.
Your one action for the next 30 minutes: list your last five CPARS ratings and, next to each, the exact document that proves each major claim. If you cannot find the source for a claim, you have just identified a gap that would have bitten you under deadline.
The one metric to track this week: the percentage of your past performance claims that have a verified source link. Start wherever you are, even if it is 30%, and push it up. When that number is high, the Thursday afternoon scramble I opened with simply does not happen. Projectory's evidence store and rubric scoring are built to hold these records and grade drafts against Section M automatically, so the reference your proposal manager needs on Monday is already assembled, tagged, and scored by the time she asks.
Frequently Asked Questions
Does agentic AI write the entire past performance volume automatically? No. It ingests, tags, scores, and drafts, but a human validates every claim and approves the final language. The gain is that people validate strong, sourced drafts instead of building them from scratch.
How does the system avoid inventing metrics? The drafting agent quotes only numbers stored in the evidence record with a verified status. If a metric is not verified, it is not written. Every claim traces back to a deliverable and a CPARS entry.
What about CUI in deliverables and client feedback? The ingestion pipeline enforces access controls, encryption, and logging consistent with NIST SP 800-171 so controlled data stays segregated and is never surfaced to the wrong audience.
How long before the evidence store pays off? The first pursuit populates it; savings compound from the second bid forward because each new proposal starts from a pre-scored base rather than a blank page.
References
- [1]Association of Proposal Management Professionals, "State of the Proposal Industry," 2024. https://www.apmp.org/page/state-of-industry
- [2]U.S. Government Accountability Office, "Contractor Performance: Actions Needed to Improve Federal Past Performance Information," 2024. https://www.gao.gov/products/gao-24-105506
- [3]General Services Administration, "Contractor Performance Assessment Reporting System (CPARS) Overview," 2025. https://www.cpars.gov/
- [4]National Institute of Standards and Technology, "SP 800-171 Rev. 3: Protecting Controlled Unclassified Information in Nonfederal Systems and Organizations," 2024. https://csrc.nist.gov/pubs/sp/800/171/r3/final
- [5]Office of Management and Budget, "M-25-21: Accelerating Federal Use of AI through Innovation, Governance, and Public Trust," 2025. https://www.whitehouse.gov/omb/information-for-agencies/memoranda/