From telemetry to judgment
Promptster derives structured, hiring-relevant evidence from a candidate’s session and organizes it into the AI Fluency rubric (ai_fluency_v1). The rubric answers how the candidate worked with AI — how they framed the task, directed the model, steered it when it drifted, and verified the result — not just what they shipped.
No part of the rubric penalizes AI usage. A candidate who leans on AI heavily but prompts well, steers deliberately, and verifies their work rates highly. Promptster measures how well a candidate works with AI, not whether they use it.
Rubric evaluation runs as a background job after the session ends. Allow a few minutes after
promptster done for signals and ratings to appear.The fluency dimensions
The rubric scores eight dimensions, grouped by the phase of work they describe. Each is evidence-backed and rated independently.Discovery
Implementation
Verification
Cross-cutting
fix_integrity requires inspecting diff content, so it is evaluated only where that content is available. When a dimension can’t be collected for a session, it is reported as not_collectible — never as a low rating.How each dimension is rated
Every dimension carries a tier (the rating) and a confidence (how much evidence backs it).
Treat
insufficient_evidence and not_collectible as absence of data, not as poor performance. They are common in short sessions or when a tool doesn’t emit a given event type.
The evidence behind the ratings
Two artifacts feed the rubric. You can read them directly to see why a dimension landed where it did.data_points_v1 — deterministic signals
Flat, countable signals computed straight from the timeline — no model judgment involved. Among them:
data_points_v1 also carries cost, testResults, and an overall confidence (high / medium / low).
fluency_signals_v1 — attributed, phase-nested signals
Richer signals nested by phase (discovery, implementation, verification) and, critically, by actor:
candidate.*— human-attributable behavior (for examplecandidate.discovery.comprehensionPromptCount,candidate.implementation.correctivePromptCount,candidate.verification.redGreenArcCount). These can raise or lower a tier.agentContext.*— model-autonomous activity (for exampleagentContext.aiDiffCount,agentContext.subagentDispatchCount). These provide context for the reviewer but never score the human — work the model did on its own is not credited to or held against the candidate.
Using signals well
- Read the evidence, not just the tier. A
developingverification_loopwithverifyIntensitynear zero tells a different story than one with highverifyIntensitybut late timing. - Respect confidence. A
strongrating atlowconfidence is a hint, not a verdict — corroborate it in the timeline. - Compare within a cohort. Ratings and signals are most meaningful relative to other candidates on the same assessment. See Cohort analysis.
