Our Methodology
We are not an opinion platform -- we are a measurement platform. This page explains exactly how, in full candor.
How Evaluation Actually Works
When you submit a work for evaluation, we build a request to an AI model (Gemini 2.5 Flash) that carries your work alongside the full text of the relevant rubric criteria — each criterion's complete description. That text is injected fresh into every single evaluation request; the model is never pretrained on the rubric bank, and it does not memorize it. The distinction matters: the model reads the criterion at the moment of judgment, exactly as it reads your work, rather than recalling it from prior training.
After comparing the two, the model returns a score for each criterion attached to evidence quoted from your own work — a line from your text, or a description of a specific element in your image. We call this evidence linking: no score stands without a citation from your own material, so you can check the judgment yourself and decide how much weight to give it.
The same mechanism underlies every deep-evaluation report on the platform, from the base evaluation to the specialized analytical reports built on the same submitted work. What differs between them is depth and scope of analysis, not the principle: reading on demand, never recalling from memory.
Confidence, Consistency, and Accuracy Limits — In Full Candor
The number shown beside every evaluation under "confidence score" is an estimate the model itself reports about how certain it is of its own judgment — not a statistically calibrated measurement validated against documented real-world outcomes. This is a distinction we insist on making explicit: a model's confidence in itself and its actual accuracy are two different things, and they do not always match.
That is why we added a second, independent signal: cross-criterion consistency. When more than one criterion measures the same facet of a skill, we calculate how closely their scores agree. High agreement means the judgment held steady across multiple angles of the work; a noticeable spread means the result deserves a closer look from you before you act on it. You will see this signal beside every evaluation where it applies.
And in full honesty, again: today's vision models — the ones that analyze images — achieve roughly 70 to 80 percent accuracy on composition and color analysis compared with the judgment of a professional human critic. That figure is not a softened marketing line; it is the known limit of this specific tool.
We surface confidence and consistency publicly for exactly this reason — not to decorate the interface, but so you know precisely when to trust a result more, and when to treat it as a starting point for discussion and review rather than a closed, final verdict.
The Rubric Bank: 232 Criteria, Eight Arts
The rubric bank today holds 232 Arabic criteria distributed across three levels for each of the eight supported arts — from scenic writing to handcraft. Each criterion describes a specific, observable skill, not a general impression.
Those three levels describe the bank's own structure: each art has its three levels, and each level carries its own set of criteria written in plain language, never left to guesswork. When your work is judged against any one of these criteria, the result is never a bare number — every score comes attached to evidence quoted from your own work explaining the basis for the judgment.
This bank was not built in one pass; it grew across successive rounds of work, and it is still growing. Every new criterion is tested against real submitted work before it is adopted, because a criterion that cannot be applied to an actual piece of work is worthless no matter how sound it looks on paper.
How Artist & AI Tools Work
The Artist & AI tools do not evaluate a work in isolation from how it was made — they evaluate the relationship between you and the tool you used. Every process starts with a disclosure: you describe, in your own words, where AI entered your work — from an initial idea you drew on, to a full execution you leaned on — and the rest of the assessment is built on that disclosure, not on an automated guess about the work's origin.
The "responsible use" assessment — and the ethics audit built on the same logic — measures how consistent what you described in your disclosure is with the tool's visible footprint in the work itself, and how much genuine artistic decision-making you kept in your own hands — choosing, cutting, rewriting — rather than handing the decision entirely to the tool. The result is not a verdict on "is AI allowed?" — it is a measure of your transparency and your own practice within its use.
The Artist-AI Compass is used as a broader diagnostic than any single work: it looks at your pattern of use over time rather than one isolated snapshot, and produces a report showing where your practice aligns with responsible disclosure and where it needs more clarity.
These tools are relatively new on the platform, and we will keep refining them as we learn more from how they are actually used. What does not change is the principle: transparency about a tool's role is part of the work itself, not an optional attachment after the fact.
How Portfolio & Identity Tools Work
The Portfolio & Identity tools shift the measurement from a single work to a body of work. "Coherence" here does not mean your works look alike — it means your choices of subject, style, and treatment show a traceable pattern over time, rather than random scatter with no connecting thread. The artist card summarizes the most notable findings from this reading in a compact, shareable form, without adding any new judgment beyond what the same evidence already supports.
This reading is built on the same kind of evidence used in single-work evaluation — evidence quoted from your actually submitted works — but compared across several pieces instead of one. The "artistic signature" report specifically describes what recurs across your works with noticeable consistency — a recurring color choice, a narrative rhythm, an angle of treatment — not a verdict on your "originality" as an artist.
The "model-committee" review is applied at the scale of an entire portfolio, serving the same purpose as the consistency signal in single-work evaluation: reducing the effect of any single judgment's bias on a reading that spans a whole body of work.
The artist statement these tools generate is built from data extracted from your portfolio analysis, not a reused template; even so, it remains a draft for you to review and edit in your own voice, not a final text attributed to you without a read-through.
How Credentials Are Issued — and What They Actually Prove
Every credential or record the platform issues — from the Art Fundamentals Certificate to more than fourteen other verifiable credential types — is built on an actual record of practice and evaluations tied to your account, not on a self-declaration or a single isolated assessment. The credential wallet and longitudinal progress reports display this record across time, not one cherry-picked moment.
"Verifiable" has a specific meaning here: every credential is linked to a digitally signed record pointing to the actual evidence it was built on — the evaluations, the works, their dates — so a third party (a school, an employer, another platform) can confirm the credential was issued from real, documented practice on the platform, not from text anyone could write themselves.
And just as clearly, here is what these credentials never claim: they do not attest to innate talent, nor to an absolute professional level, and they do not substitute for formal academic or professional accreditation outside the platform. All a credential proves is that a specific practice or achievement actually happened, against stated standards, with evidence that can be reviewed — nothing more, nothing less.
This distinction is not a marginal detail; it is the same idea the entire platform is built on, from the first evaluation to the last credential issued: we do not measure talent, we measure the skill gap and its documented progress. "You are talented" is a verdict no one can actually prove; "you practiced this skill, improved by this much, and here is your documented record" is a measurement anyone can review.