Jev classified 2,304 AI answers for $0.37, versus an estimated $58.40 with Claude Opus 5 at the same token budget. That is a 99.4% cost reduction. Only Jev was measured; the LLM figures below are pricing estimates.
Canonry 5.30 adds sentiment scores using Jev from TypeSafe. Our open-source AEO platform now shows whether answers from ChatGPT, Gemini and Claude describe a tracked brand or property favorably, alongside whether they mention it and cite its site.
The measured run: 2,304 answers, 80 seconds, $0.37
We classified every branded answer from two sweeps of roughly 200 apartment communities across three answer engines. A sweep is one run of the portfolio's tracked questions.
| Metric | Result |
|---|---|
| Apartment communities | Roughly 200 |
| Answer engines | 3 |
| Sweeps | 2 |
| Answers classified | 2,304 |
| Input tokens | 8.8 million |
| Total classifier cost | $0.37 |
| Total processing time | About 80 seconds |
| Time per request | About 0.2 seconds |
| Parallel requests | 8 |
| Throughput (calculated) | 28.8 answers/second |
Processing speed varies with answer length and account limits.
Up to 99.4% lower estimated cost than an LLM
Jev classifies each answer and returns a result with a confidence score. It does not generate a written verdict. TypeSafe charges $0.042 per million input tokens, with free output. TypeSafe pricing.
For comparison, we priced the same 8.8 million input tokens on general models, adding 250 output tokens per answer for a short JSON verdict: 576,000 output tokens in total.
| Model | Input / 1M tokens | Output / 1M tokens | Cost for 2,304 answers | Jev costs less by |
|---|---|---|---|---|
| Jev | $0.042 | Free | $0.37 measured | Baseline |
| DeepSeek V4.1 Flash, off-peak | $0.15 | $0.60 | $1.67 | 77.8% |
| DeepSeek V4.1 Flash, peak | $0.30 | $1.20 | $3.33 | 88.9% |
| GPT-5.4 nano | $0.20 | $1.25 | $2.48 | 85.1% |
| Claude Haiku 4.5 | $1.00 | $5.00 | $11.68 | 96.8% |
| Claude Sonnet 5 | $2.00 | $10.00 | $23.36 | 98.4% |
| Claude Opus 5 | $5.00 | $25.00 | $58.40 | 99.4% |
Only Jev was measured. LLM costs are estimates using uncached list prices, with no batch discount or reasoning tokens. Prompts and tokenizers differ by provider, so actual input counts will differ. Cost is input millions × input rate + output millions × output rate; percentages use unrounded figures.
Prices checked September 29, 2026: DeepSeek, OpenAI, Anthropic. These costs cover sentiment classification; answer collection and hosting are separate.
About $5 a year for this portfolio
At one sweep every two weeks, this workload becomes 26 sweeps and 29,952 branded answers a year. The measured two-sweep run therefore scales by 13.
| Model | Estimated cost / sweep | Estimated annual cost | Estimated annual batch cost |
|---|---|---|---|
| Jev | $0.18 | $4.80 | No discount assumed |
| DeepSeek V4.1 Flash, off-peak | $0.83 | $21.65 | Not included |
| DeepSeek V4.1 Flash, peak | $1.67 | $43.31 | Not included |
| GPT-5.4 nano | $1.24 | $32.24 | $16.12 |
| Claude Haiku 4.5 | $5.84 | $151.84 | $75.92 |
| Claude Sonnet 5 | $11.68 | $303.68 | $151.84 |
| Claude Opus 5 | $29.20 | $759.20 | $379.60 |
Annual figures use unrounded costs and assume unchanged volume, token use and rates. OpenAI and Anthropic offer 50% batch discounts with processing windows of up to 24 hours. For this portfolio, adding non-brand answers increases volume by about 2.4×, putting Jev at roughly $11.53 a year.
Forty seconds versus 19 hours of manual tagging
At one minute to read and label each answer, a person would spend 19.2 hours on one sweep. That is a planning assumption, not a measured human benchmark.
| Workload | Manual tagging at 1 minute / answer | Jev, extrapolated from the run |
|---|---|---|
| One sweep: 1,152 answers | 19.2 hours | About 40 seconds |
| One year: 26 sweeps | 499.2 hours | About 17.3 minutes |
The per-sweep Jev figure is half the measured 80-second run. We have not timed general LLMs on this workload. People still need to review disputed verdicts and sample the results; the table compares the initial tagging step.
Accuracy: 90.3% agreement in a small human check
Two people labeled 32 answers without seeing Jev's results. We then compared their judgments with Jev.
| Check | Result |
|---|---|
| Answers reviewed | 32 |
| Human labelers | 2 |
| Answers that expressed a stance | 31 |
| Jev verdicts agreeing with the human stance | 28 of 31 |
| Agreement rate | 90.3% |
| Favorable share: Jev | 76% |
| Favorable share: human labeler 1 | 76% |
| Favorable share: human labeler 2 | 80% |
This small check does not establish higher accuracy than an LLM. A larger human review on a separate set of answers is in progress. Sentiment is experimental, with the source quotes available for every verdict.
What the sentiment score counts
The headline is % favorable: favorable answers divided by all judged answers, shown with a 95% confidence interval and the count behind it.
| Answer or question type | How Canonry treats it |
|---|---|
| Favorable, mixed or unfavorable conclusion | Included in the judged-answer count |
| Facts only, such as address, rent or floor plans | Excluded from sentiment; does not count as unfavorable |
| No mention of the tracked subject | Excluded from sentiment; does not count as unfavorable |
| Branded question, such as “is [property] a good place to live” | Scored separately from non-brand questions |
| Non-brand question, such as “best apartments in [city]” | Scored separately from branded questions; never pooled |
| One answer comparing two tracked properties | Separate verdict for each property |
| Evidence behind a verdict | The sentence supporting it, plus the most serious complaint if reported |
Scores sit next to mention and citation coverage, per project, question and engine. The dashboard, canonry sentiment CLI, nine MCP tools and API all expose sentiment, so an agent can find poorly described properties and return the quotes behind the scores.
Turn it on for new sweeps
Sentiment is off by default. Enabling it requires both installation and project settings. Replace my-project with your project in the CLI example.
| Step | Action |
|---|---|
| Configure the installation | Add a sentiment block with a TypeSafe key to config.yaml |
| Enable the project | Use Manage sentiment, or canonry sentiment configure my-project --enabled true |
| Classify new answers | New sweeps are classified automatically after both settings are enabled |
| Classify past answers | Preview and approve classification of past answers, handling branded and non-brand questions separately |
| Read saved scores | Read stored verdicts and quotes; reading never calls Jev |
Get the source and setup instructions on Canonry's GitHub repository. Only new classifications call Jev; revisiting the evidence adds no classifier cost.