RFP Questions for Agent-Memory Systems
A general vendor security questionnaire covers infosec, SOC 2, and data handling — necessary, but silent on the questions specific to whether an agent-memory vendor's benchmark numbers mean anything. This template is that second, narrower set: twenty questions written to force a vendor to state a specific fact, not restate a marketing claim, covering ground-truth provenance, held-out evaluation, retrieval budget and granularity, model dependencies, and reproducibility. Send it alongside the standard security questionnaire, not instead of it.
What you walk away with
~15 min · 4 slots · 23 blocks- Twenty RFP questions covering ground-truth provenance, held-out evaluation, retrieval budget and granularity, model dependencies, and reproducibility.
- A "what a good answer looks like" guidance block for each section, so gaps and evasive answers are easy to spot.
- A populated document ready to send alongside the firm's standard vendor security questionnaire.
Variables
What kind of decisions or questions the agent's memory needs to answer correctly.
Live document preview
Agent-Memory RFP Questions — [VENDOR_NAME]
Issued by [FIRM_NAME] on [EFFECTIVE_DATE] for the evaluation of [VENDOR_NAME] as an agent-memory provider. Scope: [DECISION_DOMAIN]. These questions are in addition to the firm's standard vendor security questionnaire — they focus specifically on whether [VENDOR_NAME]'s memory-quality claims are checkable, not on infosec controls covered elsewhere.
Each question below is calibrated against the same structural disclosures a defensible memory benchmark makes about itself — how ground truth was produced, whether evaluation ran against genuinely unseen data, what retrieval budget was used, which model created the answers, and whether any of it is independently reproducible. A vendor who can answer all twenty specifically isn't necessarily the best performer — but a vendor who can't is asking the firm to trust a number it can't check.
1. Vendor & system profile
- Which memory architecture(s) does your system implement — verbatim/archival storage, fact extraction, temporal knowledge graph, or a combination?
- Which specific product and version is being proposed, and how long has it been in production use with other customers?
- What retrieval unit does your system return natively — a whole passage, a discrete extracted fact, a graph edge, or something else?
2. Ground-truth & labeling methodology
A vendor should be able to state plainly whether published numbers come from hand-labeled or generated ground truth, and name the actual process — not a general assurance that 'our eval team reviewed the data.'
- How was the ground truth behind your published benchmark numbers produced — hand-labeled, generated deterministically, or a mix?
- If hand-labeled, what was the review process, and who performed it?
- If an LLM was used anywhere in seeding or assisting the labeling process, how was that step's quality independently checked?
3. Held-out evaluation protocol
A defensible vendor can describe a specific, private held-out set and how its disjointness from public or training data is enforced — not just assert that their number 'wasn't overfit.'
- Do your published numbers come from a held-out set your own system's development never had access to?
- How is disjointness between the held-out set and any public or training data verified, rather than simply claimed?
- Are held-out answer keys ever published, sold, or shared with customers on request — and if so, does the held-out set still function as one?
4. Retrieval architecture & granularity
The vendor should be able to state a specific retrieval budget (k) used to produce their numbers and explain what that budget represents in their system's native retrieval unit — a vague 'we retrieve enough context' isn't a checkable answer.
- What retrieval budget (k) was used to produce your published numbers, and is it configurable in a production deployment?
- How does your system's native retrieval unit affect what a fixed budget actually returns for a single complex record?
- Will you support a matched-budget benchmark run against another system as part of our evaluation?
5. Model dependencies & configuration disclosure
A vendor should name the specific underlying model(s) their pipeline depends on and disclose whether performance has been tested across more than one — an architecture's score can swing substantially depending on which model sits inside it.
- Which underlying LLM(s) does your system depend on for extraction and/or answer generation?
- Has your pipeline been tested with more than one underlying model, and if so, how much did results vary?
- Are embeddings and model versions pinned and disclosed for any benchmark numbers you've published?
6. Reproducibility & benchmarking transparency
A vendor confident in their numbers will point to a way — a reference implementation, a documented harness, an adapter contract — for a customer to check the claim independently, not just a slide deck.
- Can a customer independently reproduce your published reference numbers, and what's required (API keys, compute, data access)?
- Is your evaluation harness or the integration contract used to produce benchmark numbers publicly documented?
- Will you support running your system through our own evaluation harness, at a matched retrieval budget against other systems under consideration?
7. Production monitoring & ongoing validation
Since production traffic has no ground-truth answer key by definition, a vendor with a real answer here describes a concrete method for monitoring answer quality on an ongoing basis — not just 'we monitor uptime and latency.'
- How do you monitor answer quality once the system is in production, where ground truth for a customer's real questions doesn't exist?
- Do you support or recommend a cross-validation method — for example, a second independent model auditing answers against only the retrieved context — for ongoing quality monitoring without ground truth?
Unfilled slots show as [VARIABLE_NAME] so the partial document still reads. Filling in the form on the left substitutes them inline.
What to do with this
Send the populated questionnaire to the vendor's sales or solutions-engineering contact alongside the firm's standard security questionnaire. Score gaps and evasive answers as data, not just missing paperwork — a vendor unwilling or unable to name a retrieval budget or a creator model is telling you something about how checkable their number actually is. File the completed questionnaire with the vendor evaluation record for the procurement decision.
FAQ
Why does this look different from our standard vendor security questionnaire?
A standard security questionnaire covers infosec, data handling, and compliance posture — necessary, but silent on whether a memory vendor's specific quality claims are checkable. This template fills that gap; it's meant to be sent alongside the standard questionnaire, not instead of it.
What if a vendor can't or won't answer some of these questions?
Treat that as informative rather than a formality to work around. A vendor unwilling to disclose the retrieval budget, the underlying model, or the labeling methodology behind their published numbers is telling you those numbers aren't independently checkable — worth weighing directly in the procurement decision.
Can we add domain-specific questions beyond these twenty?
Yes — add questions specific to {{decision_domain}} in a new section. Keep the structural sections above; they're calibrated to catch the ways a memory benchmark number can be non-comparable regardless of the buyer's specific domain.
Should we expect every vendor to have a private held-out set?
Not necessarily, but the absence of one is worth knowing explicitly. A vendor without a held-out protocol may still be a reasonable choice for a given use case — the point of asking is to know which claim you're actually evaluating, not to disqualify every vendor without one automatically.