A security questionnaire becomes useful only when procurement can translate answers into a defensible decision. Counting completed questions is not enough: a supplier can answer every field while providing little evidence, or leave a low-risk item incomplete while demonstrating strong controls where exposure is material.
This guide presents an evidence-weighted scoring model for procurement and security teams. It complements MTF Institute's broader vendor due-diligence framework by focusing on one narrower decision: how should a buyer score questionnaire responses without confusing confidence, control maturity and business risk?
Direct answer
Score a vendor security questionnaire in four layers:
- apply non-negotiable gates;
- rate the control answer;
- adjust the rating for evidence strength;
- weight the result by the exposure created by the service.
The total score supports comparison, but it never overrides a failed gate or an unresolved high-impact risk. The decision record should show both the number and the reasoning behind it.
Why percentage-complete scoring fails
A simple completion percentage treats every question as equally important and every answer as equally credible. Neither assumption is defensible.
Consider two suppliers:
- Supplier A completes 100% of the form but relies mostly on self-attestation.
- Supplier B completes 90%, provides current test evidence for critical controls and marks irrelevant questions as not applicable with an explanation.
The first supplier can look better in a spreadsheet even though the second creates more decision confidence. Procurement therefore needs separate measures for control coverage, evidence strength and exposure.
Step 1: define the exposure profile
Score the proposed relationship before scoring the supplier. A questionnaire should be proportional to what the vendor will actually do.
| Exposure dimension | Low | Medium | High |
|---|---|---|---|
| Data | Public or non-sensitive | Internal or limited personal data | Sensitive, regulated or large-scale personal data |
| Access | No system access | User-level or time-limited access | Privileged, persistent or machine access |
| Operational dependency | Easily substituted | Noticeable disruption | Material service or recovery dependency |
| Supply-chain reach | Standalone service | Limited integrations | Embedded software, infrastructure or critical subprocessors |
| Exit complexity | Simple termination | Data export or transition needed | Lock-in, migration or continuity risk |
Assign each dimension an exposure value from 1 to 3. The purpose is not to create a universal risk formula. It is to decide which questionnaire domains deserve greater weight and which failures require escalation.
This risk-based approach is consistent with the NIST Cybersecurity Framework 2.0, which is designed to help organizations understand and improve cybersecurity risk management, and with the current NIST SP 800-161 Rev. 1 update, which integrates cybersecurity supply-chain risk into organizational risk-management activities.
Step 2: establish non-negotiable gates
Some conditions should be evaluated before calculating a weighted score. A gate is a requirement whose failure changes the decision regardless of performance elsewhere.
Possible gates for a high-exposure software service include:
- named responsibility for security and incident response;
- multi-factor authentication for privileged access;
- an agreed incident-notification route and timing;
- disclosure of material subprocessors;
- tested recovery arrangements proportionate to the dependency;
- defined data return, retention and deletion terms;
- a vulnerability-handling and secure-update process.
Gates must be selected for the use case. Do not copy the strictest possible list into every procurement. A non-critical supplier handling no sensitive data should not face the same gates as an infrastructure provider with privileged access.
CISA's Secure by Design guidance is especially relevant when buying technology products: it places responsibility on producers to make security a core business requirement and highlights capabilities such as multi-factor authentication, logging and single sign-on.
Step 3: rate the control response
Use a small scale that reviewers can apply consistently:
| Control rating | Meaning |
|---|---|
| 0 | Control absent, contradicted or not answered |
| 1 | Informal or partially implemented practice |
| 2 | Defined control with an accountable owner |
| 3 | Implemented, monitored and relevant to the purchased service |
The rating should reflect the answer itself, not the confidence in the evidence. Evidence is scored separately so that a well-written assertion does not automatically become a mature control.
Step 4: apply an evidence factor
Use an evidence factor to express how strongly the submitted material supports the control rating.
| Evidence level | Example | Factor |
|---|---|---|
| Assertion | Questionnaire answer without supporting material | 0.40 |
| Documented | Current policy, procedure or architecture description | 0.70 |
| Independently tested | Relevant assurance report, test result or exercise record | 0.90 |
| Directly verified | Buyer-observed configuration, contractual commitment or scoped technical verification | 1.00 |
These factors are an MTF decision model, not an external standard. An organization should calibrate them against its risk appetite and review capacity.
For each question:
evidence-adjusted score = control rating × evidence factor
A control rated 3 but supported only by an assertion receives 3 × 0.40 = 1.20. A control rated 2 with relevant independent testing receives 2 × 0.90 = 1.80. The model makes a practical point visible: confidence in a claim matters.
Step 5: weight the domains
Assign each domain a weight based on the exposure profile. A service processing sensitive data might give greater weight to access control, encryption, privacy, incident response and deletion. A logistics dependency might place more weight on availability, recovery and concentration risk.
Example weights for a high-exposure cloud service:
| Domain | Weight |
|---|---|
| Identity and access | 20% |
| Data protection and privacy | 20% |
| Vulnerability and secure development | 15% |
| Incident response | 15% |
| Resilience and recovery | 15% |
| Subprocessors and supply chain | 10% |
| Exit and deletion | 5% |
Calculate a normalized domain score, then multiply it by the domain weight. Keep the underlying answers visible: a total of 78% can conceal one severe unresolved weakness if reviewers look only at the headline number.
A five-part decision output
The final record should contain more than approve or reject.
| Output | Decision question |
|---|---|
| Gate result | Did any non-negotiable requirement fail? |
| Weighted score | How strong is the overall evidence-adjusted control position? |
| Critical findings | Which specific risks could materially affect the organization? |
| Treatment | What must be fixed, contracted, constrained or monitored? |
| Authority | Who accepts the remaining risk and until when? |
Useful decisions include approve, approve with conditions, limited pilot, defer pending evidence and reject. Every condition should have an owner, due date and verification method.
How to handle missing and not-applicable answers
Do not treat not applicable as a perfect score. Require a short rationale and verify that the control truly falls outside the service scope. Likewise, distinguish a missing answer from missing evidence. The supplier may have a control but fail to demonstrate it; that is still a procurement risk because the buyer cannot rely on what it cannot substantiate.
Record three separate states:
- control not required for this use case;
- control claimed but evidence insufficient;
- control required and not implemented.
These states lead to different actions and should not collapse into the same zero.
Common scoring mistakes
- using one questionnaire and one weighting model for every supplier;
- allowing a high total score to override a failed critical gate;
- awarding full credit for a certification logo without checking scope and date;
- confusing policy existence with implementation;
- penalizing an explained
not applicableanswer as if it were a failed control; - failing to connect conditions to contracts, onboarding and monitoring;
- keeping the score but losing the evidence and reviewer rationale.
The management principle
Questionnaire scoring is not a contest to produce the most precise-looking percentage. It is a governance mechanism for making uncertainty explicit. A useful model shows what the supplier claims, what evidence supports the claim, how the service changes the buyer's exposure and who owns the residual decision.
Professionals building this capability can explore MTF Institute's Strategic Procurement, Sourcing & Vendor Risk Management program, which connects supplier evaluation with sourcing, contracting, risk and performance decisions.