standard equipment · governance in detail

Don't hope for accuracy. Guarantee it.

Kitewing's system is designed so you can measure accuracy, test it, and improve it. Deploy an answer bot to thousands of teammates with full confidence in the correctness of the system.

ch. 01 · golden prompts

Measure context changes apples-to-apples.

Golden prompts are your gold standard for the right answers to the most important questions. You define which questions matter to your business, and what the agent needs to know about each piece of data.

Whenever any change is made to the agentic system, your golden prompts will tell you whether the agents got smarter or dumber.

templates
versioned, admin-approved
scoring
automatic, on every run
history
golden score tracked over time
Golden investigations: approved question templates with quality scores tracked on every run

ch. 02 · audits

Track and improve accuracy over time.

Every answer, every graph, every report and dashboard is automatically audited against known-good accuracy criteria.

Know right away if an exec got a questionable answer. Quantify the accuracy of the data, track its improvement day by day, while identifying and fixing outliers.

log
every query, join and chart
filter
workspace · member · investigation
provenance
printed on every report
Audit logs: every query, source and operation on the record, filterable by workspace, member and investigation

ch. 03 · rubrics

Fully custom grading rubrics.

Answers are graded against your rubrics before they reach a reader: data quality, answer quality and visualization quality, each scored on its own scale.

You know your business best. You define what it means to be right and wrong, and get notified of any regressions before shipping new integrations or instructions.

criteria
data · answer · visualization
gate
before an answer is released
low scores
flagged for review
Rubric grading: every investigation scored for data, answer and visualization quality before release

ch. 04 · staging

Rehearse, then release.

A full staging environment for your AI system. Deploy any new context, skills or instructions to staging first, and observe qualitative and quantitative changes before shipping to everyone.

Never be surprised by a regression. Catch it in staging before you ever ship out to the team.

rehearsal
against the golden investigations
compare
live vs. staging, side by side
release
when the scores hold
Staging: changes rehearsed against the golden investigations, scores compared live versus staging before publishing

ch. 05 · role-based skills

Expertise, assigned by role.

Measurements, concepts and calculations vary by role. So should your skills.

Confidentially assign your skills by role, department or user group. Ensure each team's analysis happens the way it's supposed to happen.

content
definitions · method · pitfalls
assignment
per group — each role its playbook
editing
scoped — e.g. admins only
revisions
change-tracked — accept / reject
Role-based skills: reusable analysis playbooks the agent reads on demand, assigned per group with editing scoped to admins and every edit change-tracked for review