Our Methodology
We are not an opinion platform — we are a measurement platform. This page explains how, in full candor.
How evaluation actually works
When you upload your work, a request is built for an AI model (Gemini 2.5 Flash) that includes your work plus the full text of the specific rubric criteria — the description and all five levels for each one. This text is injected into every single evaluation request; the model is not pre-trained on the rubric bank. The distinction matters: the model "reads" the criteria at the moment of evaluation, exactly as it reads your work, then compares the two and returns a score per criterion with evidence drawn from your text or image.
What does "confidence" actually mean?
The number shown next to each evaluation is an estimate the model itself reports about how sure it is of its own judgment — not a statistically calibrated measurement against real-world data. That is why we added a second, independent signal: when more than one criterion measures the same skill, we calculate how much their scores agree with each other. High agreement means the judgment held steady across multiple angles; a notable spread means the result deserves a second look from you. You will see it next to each evaluation where it applies.
Known accuracy limits
In full honesty: current vision models (for image analysis) achieve roughly 70-80% accuracy analyzing composition and color compared to the judgment of a professional critic. We are not at "perfect" yet, and we will not claim to be. We surface confidence and consistency publicly for exactly this reason — so you know when to trust a result more, and when to treat it as a starting point for discussion rather than a final verdict.
The rubric bank: where did it come from?
200+ Arabic criteria designed across eight arts, a five-level structure built to be backed by evidence from your own work specifically. We are gradually documenting the pedagogical reference behind each criterion — you will see it displayed in the rubric bank where available.
Why this approach?
We do not measure talent. We measure the skill gap. The difference is vast: "you are talented" is a verdict; "your skill is 3 of 5, and these exercises can raise it" is a measurement. This page itself is part of that commitment: whatever we do not know for certain, we will say so plainly instead of glossing over it.