Most AI maturity assessments hand back a single number and a level name. That number is the least useful part of the exercise. This blueprint gives you a working assessment you can take right now, the exact scoring rules it uses, the design decisions that keep it honest, the fixtures it is tested against, and the prompt to build your own version for your own sector.
Try it first
Question 1 of 15. Strategy & value.
Strategy & value · 1 of 3
1 of 15
Answer for what happens now, not what is planned. Choosing not verifiable is a real answer and is never scored as zero. Everything stays in this browser tab and is submitted nowhere.
Fifteen questions, one at a time, about five minutes. Answers stay in this browser tab. Nothing is submitted, stored or tracked, there is no sign up, and your result is never held back behind a form.
Why most maturity assessments quietly fail
The standard pattern is twenty questions, a five point scale, an average, and a level name at the end. It fails for a specific and fixable reason: averaging ordinal answers across unrelated capabilities produces a number that describes nothing real. An organisation with strong governance and no measurement, and an organisation with strong measurement and no governance, can land on exactly the same middling total. The total tells them the same thing. Their actual next moves are completely different.
The second failure is the missing answer. Most assessments give you no way to say that you cannot verify something, so people guess. A guess entered as a three is indistinguishable from a verified three, and the result is presented with the same confidence. Not knowing whether pre-deployment review happens is a finding in its own right, and it usually points at a bigger problem than a low band would.
The third failure is scoring intent. Ask what your organisation is committed to and you measure enthusiasm. Ask what happened the last three times someone deployed something and you measure practice. Every question below is written about observable current behaviour, which is why several of them are uncomfortable to answer honestly.
The five capabilities this assessment looks at
Five areas, three questions each. Five is a deliberate balance for this example: enough to separate capabilities that genuinely move independently, few enough that a reader can hold the whole result in mind and act on it. A different context might justify four or six, and the scoring rules do not change if you pick a different number.
Strategy and value asks how work gets chosen, how success is defined before the work starts, and what happens when the value does not appear. Operational adoption asks whether AI use is individual habit or a repeatable process with review built into it. Governance and control asks what staff can actually rely on about permitted tools, permitted data, review before deployment and escalation when something goes wrong.
Skills and ownership asks whether training matches real work, whether expert help is reachable, and whether anyone owns the outcome rather than the rollout. Measurement and learning asks whether value is evidenced, whether anything is monitored after launch, and whether lessons change what happens next.
The bands are named Ad hoc, Emerging, Defined, Repeatable and Embedded. Each question carries a written behavioural anchor for every band, so you are choosing between descriptions of what happens, not between numbers on a scale you have to interpret.
How the scoring works, and why there is no total
Each capability is scored from its three answers, and answers marked as not verifiable are excluded from the calculation rather than treated as zero.
With three known answers the band is the median of the three. The median of three ordinal answers is one of the answers you actually gave, so the band is always a description something in your organisation genuinely matches. A mean of 2.67 describes nothing, and the moment you display it people start comparing decimals that carry no meaning.
With two known answers and one unknown, the assessment deliberately takes the lower of the two known levels and marks the capability as partial evidence. This is not a median and it is not an average of the pair. It is a conservative floor, chosen so that a missing answer can never raise a result. If you rebuild this, do not describe the rule as a two value median, because that would suggest a midpoint the engine never calculates.
With one or zero known answers no band is reported at all. The area is returned as not assessable and appears under visibility gaps instead. A low band and an invisible capability are different findings and the assessment refuses to blur them.
When the known answers for one capability span two bands or more, the result carries a mixed evidence note naming the range. That is deliberate. A capability sitting at Defined because one practice is Embedded and another is Ad hoc is an unstable Defined, and the person reading the result should know that before they plan around it.
The lowest scored capability is reported as your capability constraint. If several tie at the lowest band they are all shown, unranked, because the answers genuinely do not separate them. There is no overall score at any level of the output. That is not an omission to be fixed later; it is the reason the result is worth acting on.
What happens when two people answer differently
Run this with one person and you get one person’s visibility. Run it with three and the disagreements can be more informative than the bands. The example below is controlled fixture data, not customer responses, scored by exactly the same engine as the live assessment above.
The pattern to look for is a gap of two bands or more between roles on the same capability. In this controlled fixture that happens in one capability only, and the table names it rather than leaving you to infer it. A gap that size is not noise and it is not somebody being wrong. It usually means the policy exists and has not reached the people doing the work, or the practice exists and leadership cannot see it. Both are real problems with different fixes, and averaging the raters together destroys the evidence for either one.
Controlled example data, not customer responses
The same fifteen questions answered by three roles in one organisation. Bands are computed by the same engine as the live assessment above.
Capability
Leadership
Operations
Security
Strategy & value
Repeatable
Defined
DefinedPartial evidence
Operational adoption
Repeatable
Repeatable
DefinedPartial evidence
Governance & controlMaterial disagreement
Repeatable
Emerging
Emerging
Skills & ownership
Repeatable
Defined
Defined
Measurement & learning
Defined
Emerging
Emerging
Governance & control shows a gap of two or more bands between roles. Governance & control: Roles disagree from Emerging to Repeatable. Average the three people together and that gap disappears into a middle value that nobody in the room actually reported.
The design decisions that make it credible
One question at a time, with progress shown. Fifteen questions on one page reads as a form to be completed quickly. One at a time with the capability name visible reads as a set of questions worth thinking about, and the answers change accordingly.
Behavioural anchors instead of agreement scales. Nobody can calibrate strongly agree. A respondent can more readily verify whether pre-deployment review happened the last three times, so every option is a sentence describing an observable state.
Not verifiable offered on every question, in the same visual weight as the other options. Making it a smaller or greyed choice tells people it is a failure to answer, and they stop using it.
Horizontal band tracks rather than a radar chart. Radar charts on five ordinal axes create an area that looks quantitative and is not: rotate the axes and the shape changes while the data does not. Five stacked rows read correctly on a phone, screenshot cleanly into a slide, and cannot imply precision that the underlying answers do not support.
No numbers in the output. Bands are named, never numbered, and no decimal appears anywhere. The moment a 3.4 is on screen someone will compare it to a 3.6 and treat the difference as meaningful.
Next moves tied to specific answers. Each suggested move is attached to the question that produced it, so the recommendation is traceable back to something you said rather than generic advice appended to a level name.
The result is never held back. No email, no name, no account and no gate sits between the last question and the result. An assessment that withholds the finding until you identify yourself is a lead form wearing a methodology, and respondents answer it accordingly.
Test the assessment before you trust it
An assessment is a small piece of software making claims about an organisation, so it gets tested like software. These are the controlled inputs this page is checked against, with the outcome each one is required to produce. They are test fixtures chosen to pin behaviour at the edges. They are not benchmark data, not sector averages and not anyone’s real answers.
Every outcome shown below is computed live by the same scoring engine that runs the assessment at the top of this page, so if the engine ever changed, these rows would change with it.
Controlled test inputs, not benchmark data
Floor case
All fifteen answers at Ad hoc.
Strategy & value: Ad hoc. Operational adoption: Ad hoc. Governance & control: Ad hoc. Skills & ownership: Ad hoc. Measurement & learning: Ad hoc. Capability constraint: Strategy & value and Operational adoption and Governance & control and Skills & ownership and Measurement & learning.
The lowest anchor on every question produces the lowest band on every capability, with no rounding upward.
Ceiling case
All fifteen answers at Embedded.
Strategy & value: Embedded. Operational adoption: Embedded. Governance & control: Embedded. Skills & ownership: Embedded. Measurement & learning: Embedded. Capability constraint: Strategy & value and Operational adoption and Governance & control and Skills & ownership and Measurement & learning.
The highest anchor produces the highest band, and a tie at the lowest band across all five capabilities is reported unranked rather than picking a winner.
No visibility at all
All fifteen answers marked not verifiable.
Strategy & value: Not assessable. Operational adoption: Not assessable. Governance & control: Not assessable. Skills & ownership: Not assessable. Measurement & learning: Not assessable. No capability constraint reported.
Unknown answers are never scored as zero. Every capability returns as not assessable and no capability constraint is invented.
One unknown, two known
Strategy and value answered Emerging, not verifiable, Embedded. Everything else Defined.
With two known answers the engine takes the lower of the two as a conservative partial evidence band, so an unknown answer can never raise a result. The mixed evidence note still names the full range of what was reported.
Two unknowns in one capability
Strategy and value answered Emerging, not verifiable, not verifiable. Everything else Defined.
Strategy & value: Not assessable. Operational adoption: Defined. Governance & control: Defined. Skills & ownership: Defined. Measurement & learning: Defined. Capability constraint: Operational adoption and Governance & control and Skills & ownership and Measurement & learning.
One known answer is not enough evidence for a band. The capability moves to visibility gaps instead of being scored low.
Wide spread inside one capability
Strategy and value answered Ad hoc, Defined, Embedded. Everything else Defined.
Strategy & value: Defined, evidence spans Ad hoc to Embedded. Operational adoption: Defined. Governance & control: Defined. Skills & ownership: Defined. Measurement & learning: Defined. Capability constraint: Strategy & value and Operational adoption and Governance & control and Skills & ownership and Measurement & learning.
The median of three is one of the answers actually given, and a spread of two bands or more triggers the mixed evidence note so an unstable band is not read as a settled one.
Deliberate tie at the lowest band
Strategy and value and Operational adoption at Ad hoc. The other three at Repeatable.
Strategy & value: Ad hoc. Operational adoption: Ad hoc. Governance & control: Repeatable. Skills & ownership: Repeatable. Measurement & learning: Repeatable. Capability constraint: Strategy & value and Operational adoption.
When capabilities tie at the lowest band, both are shown as constraints, unranked, because the answers do not separate them.
Multi-rater disagreement
The three role fixtures shown earlier on this page, scored separately and compared per capability across 5 capabilities.
Governance & control: Roles disagree from Emerging to Repeatable. No other capability reaches a two band gap in this fixture.
A gap of two bands or more between roles is flagged as a material disagreement rather than averaged away.
What happens when someone submits an assessment
The live example on this page submits nowhere. Your answers stay in the browser tab, and closing it discards them.
In a published version that collects responses, the order matters more than the mechanism. Show the participant their full result first. No name, no email, no account, nothing between the last question and the finding they came for.
If you want a copy of a response, add a clearly separated optional step after the result. Label it as what it is: send my result, or share this result with the person who published it. Optional respondent name, optional work email, optional role.
State exactly what is sent. In this design that is the raw answers, the derived band and evidence state for each capability, the assessment version and the completion date. If you send anything else, say so.
If you do not need responses back, do not collect them. A local-only assessment is a complete and legitimate version of this, and it is the right default.
A response carrying a name, an email or a role is identified. Do not describe it as anonymous, and do not describe it as confidential unless you can actually keep that promise.
Data handling is part of the design
Read your own questions back before you publish. Honest answers to this assessment describe missing escalation routes, weak controls, unmonitored workflows and training that does not exist. That is a sensitive picture of an organisation, and in a multi-rater run it can also be attributable to a named individual.
Collect the minimum that makes the response useful to you. If the role is what you need for interpretation, ask for the role and leave the name out.
Tell the participant who receives the responses and what you will use them for, in plain words, next to the point where they decide whether to send.
Set an expectation about how long responses are kept and how someone can ask you to delete theirs. What is appropriate depends on your context and your obligations, and this blueprint is not legal advice, so decide it deliberately rather than by default.
Avoid free text boxes inviting incident detail unless you genuinely need them. An open field on a governance question collects exactly the material you would least like sitting in an unmanaged inbox.
If the answers should not leave the browser, run it local-only and ask people to bring their result to a conversation instead.
If you need genuine anonymity, an identified web form is not the instrument. Use a process that actually provides anonymity guarantees, and do not imply anonymity you have not built.
Design for the retest
Stamp every result with an assessment version and a completion date. Those two fields are what make a second run comparable with the first, and they cost nothing to add at the point of building.
Reassessment is the part that makes this worth running at all. Capability moves as people, tools and obligations change, and a single snapshot cannot tell you whether anything you did afterwards worked.
If you materially change a question, an anchor or a scoring rule, increment the version and stop comparing across the boundary silently. You have a different instrument, and a movement that is really a rewording is worse than no comparison at all.
Compare per capability, and treat a capability that moved out of a visibility gap as its own finding. Something becoming visible is a real change even when the band that appears is low.
Do not manufacture an overall improvement percentage. There is no total in the first result, so there is nothing legitimate to take a percentage of.
The master prompt
This builds the whole assessment as a single self contained page, including all fifteen questions and every behavioural anchor used above. It carries no Sentway dependency, so it is useful whether or not you ever publish through us. It specifies the scoring rules explicitly, because that is the part an AI will otherwise invent.
Build the full assessment
Build a single self contained HTML page: an AI adoption maturity assessment for [YOUR SECTOR OR AUDIENCE]. Plain HTML, CSS and vanilla JavaScript in one file, no build step, no external libraries, no network requests, no analytics, no storage. All answers stay in the page. This is local-only by default.
PURPOSE AND DISCLAIMER
The page is a structured self-assessment of current observable practice. Show a visible line stating that it is not an audit, not a benchmark, not a certification, not a compliance assessment and not a diagnosis. Do not imply endorsement by any organisation or framework.
CAPABILITIES
Five capability areas, three questions each, fifteen questions total: Strategy and value, Operational adoption, Governance and control, Skills and ownership, Measurement and learning.
QUESTIONS AND ANCHORS, USE EXACTLY THIS WORDING AND THIS ORDER
Present the five band options in the order given below, always Ad hoc first and Embedded last, and always offer the not verifiable option last with the same visual weight as the other five.
Strategy & value
Q. How are potential AI use cases normally chosen?
Ad hoc: Individuals or teams try things when an opportunity occurs, with no shared prioritisation.
Emerging: Teams identify use cases independently and may discuss value, but there is no consistent organisation-wide method.
Defined: Important use cases are usually considered against agreed business priorities, although the process varies between teams.
Repeatable: A repeatable review considers expected value, feasibility, risk and an accountable owner before significant work starts.
Embedded: Evidence from live AI use changes which use cases are expanded, redesigned or stopped.
Not verifiable: I don't know / I can't verify this
Q. Before a significant AI workflow goes live, how is success normally defined?
Ad hoc: It usually is not. The team mainly learns whether the tool feels useful after launch.
Emerging: There is usually a general intended benefit, such as saving time or improving service, but no specific measure.
Defined: Some important uses have an agreed success measure or baseline, while others still launch without one.
Repeatable: Significant uses normally have a named owner, a defined outcome and at least one measure before launch.
Embedded: Measures are reviewed after launch and materially influence whether the workflow continues, changes or scales.
Not verifiable: I don't know / I can't verify this
Q. What normally happens when an AI initiative is not delivering enough value?
Ad hoc: There is no regular review, so it may continue because it already exists.
Emerging: A team or manager may stop it informally if dissatisfaction becomes obvious.
Defined: Important initiatives are reviewed occasionally, but stopping criteria are not consistent.
Repeatable: There are explicit review points where evidence can lead to continuation, redesign or stopping.
Embedded: Portfolio decisions routinely redirect investment away from weak uses and toward better-performing ones.
Not verifiable: I don't know / I can't verify this
Operational adoption
Q. How is AI used in day-to-day work?
Ad hoc: Mostly as individual experiments or one-off tasks.
Emerging: Some people repeat useful personal workflows, but they depend on the individual.
Defined: Some teams have shared prompts, instructions or repeatable AI-assisted processes.
Repeatable: Important AI-assisted workflows have defined steps, owners, handoffs and review points.
Embedded: AI is embedded in repeatable operating processes that are monitored and deliberately improved over time.
Not verifiable: I don't know / I can't verify this
Q. What normally happens when one team finds an AI workflow that works well?
Ad hoc: Nothing systematic. Other people may never hear about it.
Emerging: It is shared informally with colleagues.
Defined: Reusable prompts, examples or instructions are shared within or between some teams.
Repeatable: Successful workflows can be standardised with ownership, guidance and appropriate controls.
Embedded: Successful patterns are deliberately evaluated for wider rollout, and weak patterns are retired based on evidence.
Not verifiable: I don't know / I can't verify this
Q. How is human review handled in AI-assisted work?
Ad hoc: There is no defined expectation; people decide for themselves.
Emerging: People generally know they should check important outputs, but the level of review is informal.
Defined: Common or sensitive uses have some agreed review expectations.
Repeatable: Important workflows define where human judgement, approval or intervention is required.
Embedded: The amount and type of human review is deliberately matched to risk and monitored in live operation.
Not verifiable: I don't know / I can't verify this
Governance & control
Q. What can staff reliably find out about which AI tools and data they are allowed to use?
Ad hoc: There are no organisation-wide rules or approved-tool expectations.
Emerging: There is informal guidance, but people often rely on judgement or what colleagues do.
Defined: Approved tools or basic data-handling rules exist, but gaps or inconsistent interpretation remain.
Repeatable: Approved tools, prohibited or sensitive data rules, responsibilities and exceptions are documented and accessible.
Embedded: Relevant controls are monitored or enforced, and guidance is reviewed when tools, risks or regulation change.
Not verifiable: I don't know / I can't verify this
Q. What happens before a new or higher-risk AI use is put into operation?
Ad hoc: No defined review is required.
Emerging: Someone may ask a manager, security or legal colleague informally.
Defined: Some classes of AI use receive additional review, but the trigger and process vary.
Repeatable: A defined risk-based review brings in relevant expertise such as security, privacy, legal, domain or human oversight.
Embedded: Higher-risk uses are re-reviewed when material changes, incidents or monitoring evidence indicate that controls may no longer be sufficient.
Not verifiable: I don't know / I can't verify this
Q. If an AI system produces a harmful, seriously incorrect or unexpected result, what normally happens?
Ad hoc: There is no defined route beyond the individual dealing with it.
Emerging: People generally know who they would ask for help, but handling is informal.
Defined: There is a known escalation route for at least some serious issues.
Repeatable: Incident and escalation responsibilities are documented, with significant events recorded and reviewed.
Embedded: Patterns from incidents and near-misses are reviewed and used to update controls, guidance or system design.
Not verifiable: I don't know / I can't verify this
Skills & ownership
Q. What practical AI training do people in relevant roles receive?
Ad hoc: No organised training or guidance is provided.
Emerging: Optional generic resources or informal peer learning are available.
Defined: Most relevant people receive basic guidance on using approved AI tools safely and effectively.
Repeatable: Training is tailored to roles and covers the real tools, workflows, limitations and responsibilities people face.
Embedded: Capability is refreshed over time using observed needs, incidents, changing tools and examples from real work.
Not verifiable: I don't know / I can't verify this
Q. When someone is unsure how to use AI safely or effectively, what support is available?
Ad hoc: They mainly experiment, search online or work it out themselves.
Emerging: They can usually ask experienced colleagues informally.
Defined: Named contacts, channels or guidance exist for common questions.
Repeatable: Relevant expert support and escalation are easy to access for technical, security, privacy, legal or operational questions.
Embedded: Recurring questions and support issues are analysed and used to improve guidance, training and tool design.
Not verifiable: I don't know / I can't verify this
Q. For an important AI-assisted workflow, who is accountable for the final outcome?
Ad hoc: Accountability is unclear or effectively sits with whoever happens to use the tool.
Emerging: An individual or manager is implicitly responsible, but the role is not clearly defined.
Defined: Important workflows usually have a named business or operational owner.
Repeatable: The owner has clear responsibility for quality, risk, human oversight and escalation, not just delivery.
Embedded: Accountability is maintained through the workflow lifecycle and reviewed when ownership, tools or operating conditions change.
Not verifiable: I don't know / I can't verify this
Measurement & learning
Q. How do you know whether AI is creating useful value?
Ad hoc: Mostly through anecdotes, enthusiasm or visible usage.
Emerging: Teams sometimes estimate time saved or collect informal feedback.
Defined: Some important uses have defined measures of value, quality or adoption.
Repeatable: Relevant measures are reviewed regularly against an agreed baseline or expected outcome.
Embedded: Value evidence routinely influences investment, redesign, rollout and stopping decisions.
Not verifiable: I don't know / I can't verify this
Q. How is the quality or risk of a live AI workflow monitored after launch?
Ad hoc: It is not monitored unless somebody reports a problem.
Emerging: Teams may spot-check outputs or react to complaints.
Defined: Some important workflows receive periodic checks or review.
Repeatable: Important workflows have defined monitoring, thresholds or review routines with clear ownership.
Embedded: Trends, incidents, user feedback and material changes trigger structured re-evaluation and improvement.
Not verifiable: I don't know / I can't verify this
Q. What happens to lessons from successful and failed AI work?
Ad hoc: They mostly remain with the people involved.
Emerging: Lessons are shared informally when someone remembers to do so.
Defined: Some teams document useful patterns, failures or examples for others.
Repeatable: Lessons are deliberately captured and used to change standards, training, workflow design or future decisions.
Embedded: Learning from operational evidence is continuous and visibly changes the organisation's AI practices over time.
Not verifiable: I don't know / I can't verify this
QUESTION WRITING RULES IF YOU ADAPT THEM
Every question asks about observable current behaviour, never intent, commitment, ambition or plans. Each option is a written behavioural anchor describing an observable state. Never use agreement scales. Band names in order: Ad hoc, Emerging, Defined, Repeatable, Embedded. Never display a band number.
SCORING RULES, IMPLEMENT EXACTLY
Score each capability from its three answers. Exclude every not verifiable answer from the calculation entirely. Never treat an unknown as zero and never substitute any other value for it.
Three known answers: take the median of the three.
Two known answers and one unknown: deliberately use the LOWER of the two known levels as a conservative partial-evidence band. Do not average the two. Do not describe this as a median of two, because a two-value median is not defined here and the rule is a conservative floor, not a midpoint. Mark the capability as partial evidence.
One or zero known answers: report no band. Mark the capability as not assessable and list it under a separate visibility gaps heading. Do not score it low.
If the known answers for a capability span two bands or more, add a mixed evidence note naming the range, for example evidence spans Emerging to Embedded.
Report the lowest scored capability as the capability constraint. If several tie at the lowest band, show all of them unranked.
Do NOT calculate, display or imply an overall score, a total, a percentage, a grade or an organisation-wide level. There is no aggregate output of any kind, and no decimal number appears anywhere.
OUTPUT
Show one question at a time with a progress indicator, the current capability name, and a working back button that preserves the previous answer.
On completion show five horizontal band tracks, one per capability, each labelled in text with its band name and any partial evidence or mixed evidence note. Do not use a radar chart, a spider chart or a gauge.
Then the capability constraint, then any visibility gaps, then a short list of next moves.
Each next move must be deterministically tied to a specific low or unknown answer the person actually gave, and must name the answer that produced it. No generic advice.
Stamp assessment_version and completed_on on the result.
RESULT BEFORE ANY IDENTITY REQUEST
The participant sees the full result immediately, with no name, no email and no sign up required. Never place a form, a gate or a paywall between the last question and the result.
If, and only if, the person running the assessment wants a copy, add a clearly separated optional step AFTER the result, labelled something like Send my result. It may ask for respondent name, work email and role, all optional. Beside it, state plainly who receives the response and that a response sent this way is not anonymous. Do not claim anonymity.
By default do not submit anything anywhere. Local-only mode is the default and must remain usable with the submit step removed.
ACCESSIBILITY
Use a real fieldset and legend per question with radio inputs, so the question is announced with its options.
Full keyboard operation, visible focus states, and a logical tab order.
Announce the question change with one small dedicated aria-live="polite" region carrying the question position and capability name. Do not put aria-live on the whole page or on the results container.
Colour is never the only carrier of meaning: every band is labelled in text. Contrast at least 4.5 to 1 for body text.
FIXTURE TESTS, CONFIRM ALL OF THESE BEFORE YOU CALL IT DONE
All fifteen answers Ad hoc: all five capabilities return Ad hoc.
All fifteen answers Embedded: all five capabilities return Embedded.
All fifteen answers not verifiable: all five capabilities return not assessable, and no capability constraint is reported.
One capability answered Emerging, not verifiable, Embedded: partial evidence, band Emerging as the conservative lower of the two known levels, mixed evidence spanning Emerging to Embedded.
One capability answered Emerging, not verifiable, not verifiable: not assessable, listed as a visibility gap, never scored low.
One capability answered Ad hoc, Defined, Embedded: band Defined, mixed evidence spanning Ad hoc to Embedded.
Two capabilities deliberately tied at the lowest band: both are reported as constraints, unranked.
MULTI-RATER EXTENSION, OPTIONAL
If several people complete it, show each capability as a row with one column per respondent. Flag any capability where two respondents differ by two bands or more as a material disagreement. Never average respondents together and never produce a combined organisational score.
CONSTRAINTS
No overall score at any level, in any form. No band numbers and no decimals in the interface. Never score an unknown as zero. No invented benchmarks, testimonials, statistics, logos or ratings. No claim that the result is an audit, a benchmark, a certification or a diagnosis. Mobile first, no horizontal page overflow at 390px wide.
Produces: A single HTML file with the full fifteen anchored questions, per capability scoring, conservative unknown handling, visibility gaps, named bands, traceable next moves, the fixture tests it must pass, and no overall score. The result is shown before any optional identity step.
What a structured response looks like
If you collect responses rather than leaving them in the browser, this is the shape worth storing. It keeps the raw answers alongside the derived bands, which means you can rescore later if you improve the anchors without going back to the respondents.
The assessment version matters more than it looks. Change one anchor and the results are no longer comparable with earlier responses, so the version stamp is what stops you drawing a trend line across two different instruments. The example below is derived from the controlled fixture data above, and the completion date is a placeholder rather than a date, so nothing here can be mistaken for a real submission.
Illustrative structured response shape, not a live customer submission
The same structure applied to commercial process: whether qualification is consistent, whether the pipeline reflects reality, whether CRM hygiene survives a busy month, whether coaching happens on a rhythm, and whether anything is measured beyond the number. It surfaces process constraints. It must not become a scoreboard for individual sellers, because a self-assessment that carries consequences for the person answering stops being evidence about the process.
Rebuild this assessment for sales process maturity, keeping the five capability structure, three questions per capability, the behavioural anchors, the not verifiable option and every scoring rule exactly as specified. Use these five capabilities: qualification consistency, pipeline and forecast hygiene, CRM and data discipline, coaching and enablement, and measurement and review. Write anchors about observable behaviour in a normal week. Add a visible line stating that this assesses the process, not individual people, and that it must not be used for individual performance scoring, ranking or compensation decisions.
Customer onboarding health scorecard
Onboarding fails in the joins: unclear first steps, unowned handoffs, invisible progress and risks nobody names until renewal. This variant assesses the process the provider runs, and the constraint it returns is something the provider can fix. Keep it pointed at your own process rather than at the customer, or you will produce a well structured way to blame people for not being onboarded.
Rebuild this assessment as a customer onboarding health scorecard for the team delivering onboarding, keeping the five capability structure, three questions per capability, the behavioural anchors, the not verifiable option and every scoring rule exactly as specified. Use these five capabilities: clarity of the onboarding path, ownership and accountability, handoffs between teams, progress visibility to the customer, and risk detection and escalation. Write anchors about what observably happens during a normal onboarding. Add a visible line stating that this assesses the provider’s process and is not an assessment of the customer.
AI governance readiness assessment
A narrower instrument that spends all fifteen questions inside governance instead of three: what is deployed and known about, what review happens before deployment, what monitoring runs afterwards, how incidents escalate, and who is accountable. Useful for finding the gap you would fail on. It is not an audit and it produces nothing you can present as compliance evidence.
Rebuild this assessment as an AI governance readiness self-assessment, keeping three questions per capability, the behavioural anchors, the not verifiable option and every scoring rule exactly as specified. Use these five capabilities: inventory and visibility of AI use, pre-deployment review, post-deployment monitoring, incident handling and escalation, and accountability and documented responsibilities. Write anchors about observable current practice. Add a prominent visible line stating that this is a self-assessment, that it is not a compliance audit, not a certification, not an assurance activity and not a legal determination, and that it must not be presented as evidence of conformity with any framework or regulation.
Before you send this to anyone
An assessment is read as an authority claim whether or not you intended one. These are the checks that keep it defensible.
State plainly, on the result screen itself, that this is a self-assessment of current practice and not an audit, a benchmark or a certification.
Check every anchor describes an observable behaviour. Any anchor containing committed, aims to, recognises the importance of or intends to is measuring enthusiasm and needs rewriting.
Confirm the not verifiable option appears on every question with the same visual weight as the other five.
Confirm no unknown answer is being scored as zero anywhere in your code, and that two known answers resolve to the lower of the two rather than an average.
Confirm no total, percentage, grade or overall level appears anywhere in the output.
Confirm the result appears immediately after the last question, with no email, name, account or gate in front of it.
Run the fixture tests listed on this page and check every outcome, including the tie case and the two unknown case.
Take it yourself, honestly, about your own organisation before anyone else sees it. If a question is easy to answer flatteringly, the anchors are too soft.
Have one person from a different level of the organisation take it. If their answers match yours exactly on all fifteen questions, the questions are not discriminating.
Check every next move names the answer that produced it, so nothing reads as generic advice.
Test the whole flow with a keyboard only, and at 390px wide.
If you collect responses, say beside the submit control who receives them, what is sent and that the response is not anonymous.
Publishing it, and collecting responses if you want them
The page above works on its own. Publishing puts it at a shareable URL. The result stays public and ungated either way, and response collection is an optional step that happens after the participant has already seen their own result.
Free plan
Publish my AI maturity assessment page on Sentway at a public URL. Keep the questions and the result public and ungated: every participant must see their own result immediately, with no email, name or sign up first. Then add an optional Send my result step after the result screen, with optional respondent name, work email and role, that submits the raw answers, the derived band and evidence state for each capability, the assessment version and the completion date. Next to that control, say who receives the response and that a response sent this way is not anonymous. Do not put an access gate on the assessment.
Produces: The assessment live on a public Sentway URL with the result always visible, plus an optional post-result submission your AI can read back.
Access controls are a separate decision. Restrict who can open the whole assessment only if you have a real reason to, never as a way to hold back somebody’s own result.
If you already know what your constraint is, skip the assessment and fix it. A well built assessment that confirms what you knew has cost you an afternoon and given you a chart.
Do not use this format for certification, conformity or assurance. A self-assessment produces no evidence of compliance with any framework, standard or regulation, and presenting it as though it does is a serious misrepresentation.
Do not use it for objective technical or security testing. Whether a control actually works is established by testing the control, not by asking somebody how mature it feels.
Do not use it for legal determinations. Questions about obligations, liability or regulatory status need qualified advice, not a band name.
Do not use it for high-stakes evaluation of individuals, including performance management, promotion, discipline or hiring. Self-reported bands are gameable the moment they carry consequences for the person answering, and everybody involved knows it. The same applies to allocating budget between teams on self-reported scores.
Do not use it where validated psychometric measurement is required. This instrument has no reliability testing, no validation study and no normative population behind it, and it does not claim any.
In those cases the honest alternatives are an independent audit or review, a technical test, professional advice, or a validated instrument, depending on which of them you actually need.
You also do not need Sentway for any of this. The master prompt produces a page that works anywhere you can host a file. Sentway is useful when you want it hosted at a shareable URL without setting hosting up, and optionally want responses coming back so your AI can read them in the same conversation where it built the page.
Methodology and sources
The design of this assessment draws on published work in two areas: how capability maturity is described and reassessed, and how survey questions and response options should be written. The sources below informed the framing. They are listed so you can check the reasoning rather than take it on trust. Methodology reviewed August 2026.
This Sentway assessment is independent. It is not an implementation of, validation of, certification against, compliance assessment for, or endorsement by any of the organisations or publications listed below, and none of them have reviewed it. The questions, anchors and scoring rules are ours, and the scoring rules are written out in full on this page so the instrument can be judged on its own terms.
Because averaging across unrelated capabilities produces a number that describes nothing real. Strong governance with no measurement and strong measurement with no governance average to the same total, while the two organisations need completely different next moves. The per capability bands and the constraint are the actionable output.
How long does the assessment take?
Fifteen questions, one at a time, typically about five minutes if you answer for what happens now rather than what is planned.
Do I have to give my email to see my result?
No, and you should not have to on any version of this. The assessment on this page shows your result immediately, and the publishing instructions in this blueprint keep the questions and the result public and ungated. If a published version collects responses, that is a separate optional step after you have already seen your result.
Is my data stored or sent anywhere?
No. The assessment on this page runs entirely in your browser, nothing is submitted, nothing is stored, and there is no sign up to see your result.
What does it mean when a capability comes back as not assessable?
It means two or more of the three questions in that area were marked as not verifiable, so there is no honest basis for a band. It is reported as a visibility gap rather than scored low, because not being able to see a capability is a different finding from that capability being weak.
Why is the median used instead of the average?
The median of three ordinal answers is always one of the answers you gave, so the band describes something that genuinely matches your organisation. An average produces decimals that carry no meaning and invite comparisons that the underlying scale does not support.
What happens when one of the three answers is not verifiable?
The unknown answer is excluded and the capability is scored from the two that remain, using the lower of the two as a conservative partial evidence band. It is not a median of two and it is not an average, and the rule exists so that a missing answer can never raise a result.
Can I use this assessment for my own clients?
Yes. The master prompt above produces your own self contained version, including all fifteen questions and anchors, and the scoring rules are written out in full so you can rebuild them in whatever you already use.
Is this an audit or a benchmark?
No. It is a self-assessment of current practice. It has no external validation, no comparison population and no certification value, and the result screen says so.
Do I need Sentway to build one of these?
No. The page the master prompt produces works anywhere. Sentway is for hosting it at a shareable URL and, if you want them, collecting responses your AI can read back.