How Mizan works

How Mizan works

Three lanes: code, model and people. The clause moves between them, so you can see the model never works alone: code checks whatever it returns, and people come in only at the end. Tap any station for its details; below, the same path step by step on a real clause.

CodeModelHuman2,719 decision unitsNot a clause: skippedQuote not verbatim: dropped, reviewContractFile or pasted textMasks10 kinds of numbersChunksVerbatim partsClassifiesClause or notRetrievesBM25 keyword searchJudgesVerdict, quote, roleVerifiesThe code gateOrdersThe review fileCommitteeApprove, reject, referNeither the file nor its text is stored; only masked parts are sent.
Code

Verifies

In
The proposed verdict
Out
An accepted verdict, or one lowered to review or abstain
Guarantee
Drops any quote that is not verbatim or not a candidate. No «non-compliant» without an explicit prohibition or condition. An amended decision sends the clause to review.

The lit station is where the clause is now. Tap any station for its details.

A station between two lanes is done by both, and its lean shows who does more: the model chunks and classifies while code checks, and code searches after a short model call.

One real clause, step by step

Fully automated from upload to review file. No human inside the review.

Follow one clause of the demonstration contract through the engine. Everything on this journey comes from a real run of the code deployed on this site, not an illustration.

Step 1 of 7

Masks

What happens

Before anything reaches the model, code replaces ID, IBAN, phone, card and other numbers with tags.

Why it matters

The model does not need customer data to judge a clause's wording. What is never sent cannot leak.

Before

يلتزم العميل حامل الهوية رقم 1023456789 وجواله 0551234567 بسداد الأقساط إلى الآيبان SA4420000001234567891234.

After

يلتزم العميل حامل الهوية رقم [هوية] وجواله [جوال] بسداد الأقساط إلى الآيبان [آيبان].

An illustrative line, run through the same masking function every contract goes through.

Arrow keys move between steps

Fully automated, and the committee decides

The review waits for no one: from upload to review file the engine runs alone, the same steps every time. People come in only at the end: the committee decides on a finished file instead of reading the contract from scratch.

Mizan, no human

  1. Masks
  2. Chunks
  3. Classifies
  4. Retrieves
  5. Judges

Review file with evidence

Ordered by priority; each clause with its decision text and reason.

Sharia committee

  • Approve finding
  • Reject
  • Refer for study

Mizan never issues a fatwa or an approval. It shortens the road to the decision.

Why trust the result?

  • It cites only what it retrieved

    A citation to a unit outside the candidate list is dropped in code, so the model cannot bring a decision from memory.

  • The quote is verbatim, or it is dropped

    Code matches the quote against the decision text after normalising diacritics, tatweel and hamza, then shows the reader the decision's own text, not the model's copy.

  • No verdict without a rule

    «Non-compliant» needs a unit with an explicit prohibition or condition, «compliant» needs one or an explicit permission, and the scope must cover the clause. Otherwise the clause is escalated, or Mizan abstains when it finds no basis.

  • Amended decisions are caught

    If a later decision amended the cited one, the clause is flagged and escalated instead of judged.

  • And it is measured

    On 70 synthetic clauses: 93% of non-compliant caught at 96% precision, against 68% and 48% for keyword search. On 4 whole real contracts (21 clauses): 71% agreement with the key. Caveats on the main page.

The technology, briefly

Where are the decisions?
1,217 decisions, split into 2,719 units, in files on the server (13 MB). They load into memory at start-up, so no database is needed.
How does it find the decision?
Keyword search (BM25) over two copies of the decisions: the text as is, and the text with contract-style sentences we added to each decision. Before searching, the model adds the clause's fiqh terms. About 17 candidate decisions come out.
Is it RAG?
Yes: we retrieve decisions, then the model judges from them alone. After it comes a code gate that checks every quote is verbatim, or drops it and sends the clause to review.
Did you train a model?
No. The knowledge lives in the decisions and the code rules, so the model changes with one setting.

We compared search methods: of 90 clauses, how many lost their governing decision?

  • Keyword search (in use now)6
  • Semantic search alone Qwen3 Embedding9
  • Semantic search alone BGE-M320
  • Keyword + semantic hybrid4
  • Keyword + reranking BGE reranker8

Lower is better. 90 clauses from the test set, 17 candidates per clause for every method. The models ran on our own machine, so no decision text left it.

The result: semantic search alone is weaker, and reranking did worse than today's search (8 against 6); we think that is because the judging model already reads every candidate and picks for itself. Keyword plus semantic is better by two clauses out of 90, so it goes into the bank version with a local model.

Isn't a general chat model enough?

A fair question, so we tested it. We gave the same ten clauses to general models (Claude Sonnet 5.5, Opus 5.5 and GPT-6 Astra) through their APIs, in 33 runs.

Without the committee's decisions

Asked to abstain without a basis, they abstained in 150 of 150 answers. Without that request they gave 56 verdicts with no citation at all, some against the committee's decision: one judged «compliant» a clause that decision ق7/ثالثاً/3 forbids, in two runs of five.

With all decisions in its context

It came close: 8 correct verdicts out of 10 on average, 32 citations all verbatim, nothing invented in our sample. But it gave the same clause different verdicts from one run to the next.

So the difference is not that a general model «doesn't know». The difference is the guarantee: a code gate that rejects any non-verbatim quote, a verdict conditioned on an explicit rule, the same steps every time, a record per clause, and numbers measured against a baseline. Sharia review cannot settle for «usually».

A small experiment: 10 clauses, 33 runs, on 1 October 2026, through the APIs rather than the chat apps.

Before launch: a local model inside the bank

Before Mizan reaches any bank, we move it to an open-source model running on the bank's own servers. The contract text then stays inside the bank's network and is not sent to any outside party.

  1. Today

    A working demo

    Runs on a cloud model, to prove the idea and measure it on the test set.

  2. Before launch

    A local model inside the bank

    An open-source model chosen by the bank, on its own servers. The contract text stays inside its network.

  3. Before go-live

    Measured first

    We run the same test set on the chosen model and show its numbers before any real use.

Why is this easy? Because the model is a single setting in Mizan, not part of its code. The gate and abstention work the same with any model.

Examples of open-source models that can run inside the bank

  • Qwen/Qwen3.8-27BApache-2.0
  • zai-org/GLM-5.3-FlashMIT
  • humain-ai/ALLaM-7B-Instruct-preview

And we tried it: Mizan ran on a laptop with Qwen3-14B and no internet. On 4 clauses it matched the expected verdict on 3, and every quote was verbatim. That is why the chosen model is measured before go-live.

Try it on the demonstration contract