Built at VIT Chennai · Gemma × GDG × OSW

PRISM

AI Bias Observatory for India

One prompt goes in. Six demographic variants come out. We measure where the model changes its mind — and hand you the receipts.

Live pipeline

streaming
  1. Prompt

    done

    sanitised + stored

  2. Gemma baseline

    done

    gemma-4-31b-it generation

  3. Variant generator

    done

    6 demographic rewrites

  4. Semantic divergence

    done

    lexical delta vs baseline

  5. Behavioural analysis

    running

    refusal + tone drift

  6. Threat score

    queued

    TES compositor

  7. Leaderboard

    queued

    ranked + archived

Divergence by axis

averaged across recent runs

0

variants evaluated

0

prompts submitted

0.0

mean threat score

0

critical findings

Bias isn't a vibe. It's a measurement.

Four instruments, one observatory. Everything you see runs on the same evaluation trace.

Variant Generator

Gender, region, caste, religion, language and income rewrites of your prompt — semantically identical, demographically different.

Semantic Divergence

Embedding-space distance between the neutral baseline and each variant, normalised against the model's own noise floor.

Threat Score

A single 0–100 composite you can argue about, backed by a per-layer breakdown you can't.

India Heatmap

Where the failures cluster geographically, aggregated from every public submission on the board.

how it works

Seven steps, one live trace.

The whole trace is public. If you don't trust the score, read the pipeline.

  1. Prompt

    done

    sanitised + stored

  2. Gemma baseline

    done

    gemma-4-31b-it generation

  3. Variant generator

    done

    6 demographic rewrites

  4. Semantic divergence

    done

    lexical delta vs baseline

  5. Behavioural analysis

    done

    refusal + tone drift

  6. Threat score

    done

    TES compositor

  7. Leaderboard

    running

    ranked + archived

Signal density

hover a hotspot

no signals recorded yet

no datamoderateelevatedcritical

meet TES

The Tri-Layer Evaluation Stack

A single judge model is a single point of bias. So we stacked three orthogonal layers and made the last one blind.

01

Semantic Layer

Every demographic variant is compared against the neutral baseline in code. We measure drift, not vibes.

lexical Δ

02

Behavioural Layer

Refusals, hedging, harshness and length are diffed across variants — where bias hides when the words look polite.

refusal ratio

03

Adversarial Layer

Gemma critiques its own outputs under a blind rubric, so the judge never sees which demographic it is scoring.

blind rubric