Skip to main content

From telemetry to judgment

Promptster derives structured, hiring-relevant evidence from a candidate’s session and organizes it into the AI Fluency rubric (ai_fluency_v1). The rubric answers how the candidate worked with AI — how they framed the task, directed the model, steered it when it drifted, and verified the result — not just what they shipped.
No part of the rubric penalizes AI usage. A candidate who leans on AI heavily but prompts well, steers deliberately, and verifies their work rates highly. Promptster measures how well a candidate works with AI, not whether they use it.
Rubric evaluation runs as a background job after the session ends. Allow a few minutes after promptster done for signals and ratings to appear.

The fluency dimensions

The rubric scores eight dimensions, grouped by the phase of work they describe. Each is evidence-backed and rated independently.

Discovery

Implementation

Verification

Cross-cutting

fix_integrity requires inspecting diff content, so it is evaluated only where that content is available. When a dimension can’t be collected for a session, it is reported as not_collectible — never as a low rating.

How each dimension is rated

Every dimension carries a tier (the rating) and a confidence (how much evidence backs it). Treat insufficient_evidence and not_collectible as absence of data, not as poor performance. They are common in short sessions or when a tool doesn’t emit a given event type.

The evidence behind the ratings

Two artifacts feed the rubric. You can read them directly to see why a dimension landed where it did.

data_points_v1 — deterministic signals

Flat, countable signals computed straight from the timeline — no model judgment involved. Among them: data_points_v1 also carries cost, testResults, and an overall confidence (high / medium / low).

fluency_signals_v1 — attributed, phase-nested signals

Richer signals nested by phase (discovery, implementation, verification) and, critically, by actor:
  • candidate.* — human-attributable behavior (for example candidate.discovery.comprehensionPromptCount, candidate.implementation.correctivePromptCount, candidate.verification.redGreenArcCount). These can raise or lower a tier.
  • agentContext.* — model-autonomous activity (for example agentContext.aiDiffCount, agentContext.subagentDispatchCount). These provide context for the reviewer but never score the human — work the model did on its own is not credited to or held against the candidate.
This separation is what lets the rubric evaluate the candidate’s judgment rather than the model’s throughput.

Using signals well

  • Read the evidence, not just the tier. A developing verification_loop with verifyIntensity near zero tells a different story than one with high verifyIntensity but late timing.
  • Respect confidence. A strong rating at low confidence is a hint, not a verdict — corroborate it in the timeline.
  • Compare within a cohort. Ratings and signals are most meaningful relative to other candidates on the same assessment. See Cohort analysis.