TypeSafe Jev · policy calls at machine speed

AI content moderation — allow, flag, or remove user content with a confidence you can defend

Paste a comment, post, or review and Jev returns whether it breaks policy, which category it falls under, and the warranted action — a typed judgment with calibrated confidence, not a black-box verdict.

Run content moderation now

This tool is pre-configured for content moderation. Paste your text and Jev returns the typed decision below — the same call the showcase example was captured from.

0 chars
Your typed decision appears hereCategory · yes/no · score, each with confidence
Typed decision + calibrated confidence — nothing to parse, nothing to hallucinate.

A real content moderation decision

Captured live from this exact tool — the text that was pasted, the typed decision Jev returned, and the measured latency and cost. Nothing here is mocked up.

Content moderation199 ms

honestly this whole community is trash and everyone posting here is an idiot who should just quit

Is ViolationYes · 92%
CategoryHarassment96%
Harassment97%
Hate2%
Safe1%
Nsfw0%
ActionRemove81%
Remove88%
Flag12%
Allow0%
typesafe/jev-1.13-20260917$0.000020

How content moderation works

01

Paste the user-generated content

A comment, a forum post, a product review — the exact text a user submitted. Jev works on the raw string, so there is nothing to pre-clean.

02

Jev scores violation, category, and action

One pass answers three typed questions: does it violate a typical community policy, which category (spam / harassment / hate / nsfw / safe) fits, and what action (allow / flag / remove) is warranted.

03

Route by confidence, not by gut

High-confidence removes can auto-action; borderline confidence lands in a human queue. Because every answer carries a probability, your review load shrinks to the genuinely ambiguous cases.

Content moderation — frequently asked

How is this more accurate than a banned-word filter?

A word list flags "you idiot" but misses a coordinated harassment campaign written politely, and it false-positives on medical or reclaimed language. Jev judges the whole message against a policy category and reports how sure it is, so subtle harassment gets caught and clinical text does not get nuked.

Can I calibrate the remove threshold to cut false positives?

Yes — the action question returns per-option probabilities. Auto-remove only above, say, 0.9 and send everything between 0.5 and 0.9 to a moderator. You trade recall against reviewer load explicitly instead of guessing.

Which policy categories are covered?

The preset covers safe, spam, harassment, hate, and nsfw — the categories most community guidelines actually enforce. They are defined as typed criteria, so you can align them to your own policy wording.

Will it hallucinate a violation that is not there?

Jev never free-writes a rationale it could invent; it returns a constrained category and a probability. If the content is clean, the safe option wins with high confidence and the violation probability stays low — there is no narrative to fabricate.

Is it fast enough for a live comment stream?

The moderation sample on this page returned in 199 ms. That is comfortably fast enough to gate a comment before it renders, rather than moderating after the fact.

Moderate the ambiguous, auto-handle the obvious

Paste a real comment and let Jev tell you — with a defensible confidence — whether to leave it, flag it, or take it down.

Moderate a sample