Skip to content

fix(eql): constant operation count for the ORE block comparator - #1148

Draft
coderdan wants to merge 3 commits into
mainfrom
fix/eql-ore-block-256-constant-op-count
Draft

coderdan wants to merge 3 commits into
mainfrom
fix/eql-ore-block-256-constant-op-count

Conversation

@coderdan

@coderdan coderdan commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Summary

The ORE block comparator (eql_v3_internal.compare_ore_block_256_term(s), behind every ordered ORE operator and the ORE operator class) took time that depended on where its operands first differ. A party who can time queries but cannot read the stored ciphertexts, such as an application user, could learn how long a prefix their query value shares with stored values. That is the information the constant-time comparator in ore.rs keeps from them. This PR makes the PL/pgSQL comparator run a constant number of operations whatever its operands, and adds a timing harness that measures it.

PL/pgSQL is not a constant-time environment. The target here is a constant operation count: the same statements and function calls on every path. C-level effects inside those statements remain (memcmp stopping early within 16 bytes, the offset a substr reads at), orders of magnitude below one statement.

What was wrong

  • Text (multi-term) values. The array comparator stopped at the first unequal term. A comparison took about 11 µs when the first of six terms differs and 47 µs when only the sixth does (t below −1500). This is the large one.
  • Single-term values. Each block after the first difference ran extra statements, and the OR short-circuit skipped a comparison on each of those blocks. On the machine measured the two nearly cancel (t = +4.9 and +13.9). With the short-circuit removed, the extra statements alone gave t = +388, about 0.7 µs per comparison. Nothing guarantees the cancellation on other hardware or Postgres builds.

Changes

  • packages/eql/src/v3/sem/ore_block_256/functions.sql
    • Term comparator: one statement per block, the same for every block. Both comparisons are always evaluated (integer |), and the result is collected into a bitmask of differing blocks. The first difference is the mask's lowest set bit, read with arithmetic (mask & -mask, then an exact base-2 log, checked against all 16,383 masks). Exactly one encrypt() and one get_bit() run on every path, equal terms included. The result is computed without a branch on the data.
    • Array comparator: single-term values keep a fast path. For multi-term values, every paired term is scanned, the first unequal term and block are latched without branching, and the order is decided with one hash. The multi-term variables are declared in an inner block, so the single-term (operator-class hot) path does not initialise them. Term counts and NULL elements are public structure and still branch.
  • Tests: a new creds-free test pins the multi-term structure (equal prefixes, NULL elements, the later-term length check); a stale comment in the length-guard sweep is updated.
  • packages/eql/tests/timing/ore_block_256/: the timing harness (see below).
  • Changeset: @cipherstash/eql patch.

Results

PostgreSQL 17.11, aarch64 (Docker on an Apple M1 Max), through compare_ore_block_256_terms. 200,000 single-term samples of 16 comparisons and 50,000 six-term samples of 8, two runs each. The classes are operands that first differ early against operands that first differ late; |t| < 5 means no detectable difference.

single term: t six terms: t cost
before +4.9, +13.9 −1552, −2575 11–47 µs per six-term comparison, depending on the data
after +1.8, +1.1 +4.5, +2.7 about 3–8% more per single-term comparison; about 31 µs per six-term comparison, always

Behaviour change

Multi-term values now length-check every paired term, so a malformed term after the first unequal one raises an error instead of being skipped. Results are otherwise unchanged.

Verification

  • All 1,500 harness checks (1–6 terms, every shared-prefix length) match plaintext order and are antisymmetric, against real ore-rs ciphertexts.
  • 11 structural edge cases (NULL elements, equal prefixes of different lengths, empty and NULL arrays) give exactly the same results as the previous implementation.
  • The 9 creds-free comparator tests pass against the EQL test database: the length guards, the NULL ordering, the array base cases, and the new multi-term structure test.
  • mise run build, codegen:parity, test:self_contained_v3, docs:validate and cargo fmt --check all pass.
  • Not run locally: the fixture-based ordering tests need CipherStash credentials, so CI is the first run of those.

Timing harness

DATABASE_URL=... packages/eql/tests/timing/ore_block_256/run.sh runs by hand against any database with EQL installed. It needs no CipherStash credentials: a detached ctgen crate encrypts real legacy ORE ciphertexts with ore-rs under a throwaway key. It refuses to time a comparator that fails its correctness checks. It applies three lessons from timing work in ore.rs, each of which produced a false signal there:

  • both classes share one randomly ordered table;
  • each sample times several comparisons;
  • it reports the plain Welch t, not a percentile-cropped maximum.

Against the previous comparator it reports six-term t of −460 to −2575.

It is not wired into CI, because timing on shared runners is too noisy to gate on.

Closes #1149

compare_ore_block_256_term(s), behind every ordered ORE operator and the
ORE operator class, took time that depended on where its operands first
differ. A party able to time queries, without reading the stored
ciphertexts, could learn how long a prefix their query value shares with
stored values: the information the constant-time comparator in ore.rs
keeps from them.

- Within a term, every block after the first difference ran extra
  statements, and the `OR` short-circuit skipped a comparison on each of
  them. On PostgreSQL 17 (aarch64) the two nearly cancel (t = +4.9 and
  +13.9 through the operator path); with the short-circuit removed the
  extra statements alone gave t = +388, about 0.7 µs per comparison.
- Across terms (text, up to six), the comparison stopped at the first
  unequal term: about 11 µs when the first term differs, 47 µs when only
  the sixth does (t below -1500).

The term comparator now runs the same statements for every block: both
comparisons always evaluated (integer `|`), one statement per block
collecting a bitmask of differing blocks, the first difference read off
its lowest set bit with arithmetic, and exactly one encrypt() and one
get_bit() on every path. The array comparator keeps a fast path for
single-term values and, for multi-term values, scans every paired term,
latches the first unequal term and block without branching, and hashes
once; its multi-term variables live in an inner block so the single-term
path does not initialise them. Term counts and NULL elements are public
structure and still branch.

Measured with tests/timing/ore_block_256 (next commit): single-term
t = +1.8 and +1.1 at about 3-8% more time; six-term t = +4.5 and +2.7 at
about 31 µs either way. Results are unchanged: all 1,500 harness checks
(1-6 terms, every shared-prefix length) match plaintext order and are
antisymmetric, and the structural edge cases (NULL elements, equal
prefixes of different lengths, empty and NULL arrays) match the previous
implementation exactly.

Behaviour change: multi-term values length-check every paired term, so a
malformed term after the first unequal one now raises instead of being
skipped. A new creds-free test pins that and the multi-term structure.
tests/timing/ore_block_256 runs by hand against a database with EQL
installed (run.sh, DATABASE_URL). ctgen encrypts real legacy ORE
ciphertexts with ore-rs under a throwaway key, so no CipherStash
credentials are needed; it is a detached crate with its own workspace.

The harness first checks the comparator against plaintext order on 1,500
pairs of 1-6 terms, then times two classes (operands that first differ
early vs late) for single-term and six-term values through
compare_ore_block_256_terms. Both classes share one randomly ordered
table, each sample times several comparisons, and it reports the plain
Welch t: separate per-class storage and percentile-cropped statistics
each produced false signals in the ore.rs timing work.

Against the previous comparator it reports t of -460 to -2575 for
six-term values; against the current one every |t| is below 5.
@changeset-bot

changeset-bot Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: cdfd573

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 12 packages
Name Type
@cipherstash/eql Patch
stash Patch
@cipherstash/stack-prisma Patch
@cipherstash/basic-example Patch
@cipherstash/e2e Patch
@cipherstash/prisma-example Patch
@cipherstash/stack-drizzle Patch
@cipherstash/stack-supabase Patch
@cipherstash/stack Patch
@cipherstash/wizard Patch
@cipherstash/bench Patch
@cipherstash/test-kit Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@coderabbitai

coderabbitai Bot commented Oct 10, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Multi-term equality still changes the operation count, and fixed-width masks break otherwise valid wider terms.

3 open findings
What changed in this PR

Makes the EQL ORE comparator’s operation count independent of the first differing ciphertext block or term and adds a manual timing harness.

Changes:

  • Reworks single- and multi-term comparator logic.
  • Adds correctness coverage and timing infrastructure.
  • Adds an EQL patch changeset.
File Description
.changeset/​eql-ore-block-256-constant-op-count.md Documents the security fix.
packages/​eql/​src/​v3/​sem/​ore_block_256/​functions.sql Implements constant-operation comparison.
packages/​eql/​tests/​sqlx/​tests/​ore_block_comparator_tests.rs Updates comparator expectations.
packages/​eql/​tests/​sqlx/​tests/​encrypted_domain/​family/​sem.rs Tests multi-term structure and validation.
packages/​eql/​tests/​timing/​ore_block_256/​README.md Documents timing methodology.
packages/​eql/​tests/​timing/​ore_block_256/​run.sh Runs correctness and timing measurements.
packages/​eql/​tests/​timing/​ore_block_256/​harness.sql Implements sampling and Welch analysis.
packages/​eql/​tests/​timing/​ore_block_256/​ctgen/​src/​main.rs Generates ORE ciphertext samples.
packages/​eql/​tests/​timing/​ore_block_256/​ctgen/​Cargo.toml Configures the detached generator crate.
packages/​eql/​tests/​timing/​ore_block_256/​ctgen/​Cargo.lock Locks generator dependencies.
packages/​eql/​tests/​timing/​ore_block_256/​ctgen/​.gitignore Excludes generator build output.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +351 to +354
-- Every paired term equal: the shorter array sorts first.
IF still_eq = 1 THEN
RETURN sign(cardinality(a) - cardinality(b))::integer;
END IF;
Comment on lines +128 to +129
--! with arithmetic rather than an `IF`, and exactly one `encrypt()` and one
--! `get_bit()` run on every path, including for equal terms. An earlier form
Comment on lines +27 to +29
-- dudect-style sampling. Each sample picks a class at random and times
-- `per_sample` comparisons of distinct random pairs of that class through
-- compare_ore_block_256_terms (the operator-class path). clock_timestamp()
@coderdan

Copy link
Copy Markdown
Contributor Author

Code review (Claude Code, /code-review xhigh, head 379c161)

Posted for the record. The main findings were checked against a scratch PostgreSQL 17 container with the old and new comparators installed side by side.

  • Correctness holds. On 20,000 random arrays (1–4 terms, with NULL elements and shared prefixes) the new comparator returned exactly what the old one did.
  • The multi-term path still leaks in two places (findings 1 and 3), and wide terms break (finding 2).
  • A cheaper multi-term form compares each term in one statement and scans blocks only on the first unequal term: about 9 µs against 26–29 µs per six-term comparison.

Findings

  1. functions.sql:352: different term counts leak a shared prefix. The still_eq = 1 early return skips encrypt()/get_bit(), so when the arrays have different term counts, timing shows whether every paired term is equal. Measured: [x] vs [x, y] ≈ 7 µs, [x] vs [z, y] ≈ 9 µs; both return −1. The timing harness misses it because its six-term classes always have equal term counts. Fix: run the hash on every path and select the result arithmetically.
  2. functions.sql:217: terms of 32 or more blocks break. diff_mask is int4 with no limit on N. With N = 32 differing only in block 31: ERROR: integer out of range (the old code returned ±1). With N = 33 differing only in block 32: 1 << 32 = 1, so the wrong block is hashed and the order is silently wrong. The same code is at lines 206 and 333/340.
  3. functions.sql:357: the NULL early return leaks. Whether the first unequal pair is the NULL pair depends on whether the earlier terms were equal. [x, NULL] vs [x, y] ≈ 7.5 µs; [x, NULL] vs [z, y] ≈ 9.8 µs. The NULL positions are identical, but timing shows whether x = z.
  4. functions.sql:331: multi-term cost. Scanning every block of every paired term (6 × 8 = 48 block statements) costs 26–31 µs per six-term comparison, about 3× the old best case, on every text ORDER BY, index probe and min/max. A prototype that compares each whole term in one statement, latches the first unequal term, then scans blocks and hashes once took ≈ 9.0 µs whether term 1 or term 6 differs.
  5. functions.sql:217: round(ln(mask & -mask)/ln(2)). glibc's log() returns early for exactly 1.0 (first block differs, or equal terms) and takes the table path for 2^k with k ≥ 1: a C-level branch on "first block vs later block". It is also the cause of the 31-block limit. A one-statement arithmetic latch per block avoids both: latch := latch + (latch = 0)::int * differs * (block + 1).
  6. functions.sql:319: duplication. The multi-term block repeats the length checks, block scan, lowest-set-bit maths and hash expression from compare_ore_block_256_term almost word for word (319–341 and 361–376 against 185–231). Fixes to the constant-time logic would have to be made twice.
  7. sem.rs:125: no CI test covers the multi-term ordering logic. The new creds-free test covers only the equal, NULL and raise branches. An off-by-one in the first_term or first_block latch would pass every creds-free test.
  8. ctgen/src/main.rs:30: harness coverage. The timing classes cover only equal-cardinality, NULL-free, 8-block inputs, so the harness cannot detect findings 1–3, yet reports |t| < 5 as clean.
  9. functions.sql:322: behaviour change. Every paired term is now length-checked, so a malformed term after the first difference raises instead of being skipped. This ships as a patch changeset; the changeset mentions it.
  10. functions.sql:314: repeated subscripts. Each iteration evaluates a[t]/b[t] about seven times. Assign ab/bb once and test the locals.
  11. run.sh:27: DROP SCHEMA IF EXISTS ore_timing CASCADE runs on whatever database DATABASE_URL names, at start and in the EXIT trap. It would destroy an existing ore_timing schema. Use a unique per-run schema name, or refuse if it exists.
  12. sem.rs:75: stale comments. The module header and the T4 doc still describe the recursion this PR removes; T4b is missing from the module's T2–T5 index.
  13. sem.rs:126: synthetic terms. The new test builds ROW(repeat('a', 65)::bytea) terms by hand. packages/eql/AGENTS.md says tests must use real ciphertexts, never synthetic blobs (existing structural tests do the same), and the file already has a term() helper.
  14. .github/dependabot.yml:215: the new detached ctgen workspace is not listed, so its locked dependencies never get Dependabot updates.
  15. functions.sql:257: two wrong doc claims. "Exactly one encrypt() and one get_bit()": the code calls get_bit twice, plus get_byte. "Every paired term is length-checked": a term paired with a NULL element is not checked, so a malformed term there passes.

Plan

  • Stage A, the comparator: findings 1–6, 10 and 15.
  • Stage B, tests and harness: findings 7, 8, 11, 12 and 14. For findings 7 and 13, a creds-free test will use fixed ciphertexts generated by ore-rs.
  • Finding 9: stays a patch; it affects only corrupt data.

Review of the constant-operation-count comparator found three remaining
data-dependent paths and a width limit:

- When paired terms were all equal and the arrays had different term
  counts, the array comparator returned without hashing, so timing showed
  whether one value's terms are a prefix of the other's (about 7 µs
  against 9 µs).
- When the first unequal pair had a NULL element, it also returned without
  hashing, so timing showed whether the terms before the NULL were equal.
- The first differing block was read off a bitmask with
  round(ln(mask & -mask) / ln(2)). glibc's log() returns early for exactly
  1.0, which is the first-block case, and the int4 mask failed at 32 or
  more blocks: "integer out of range" at 32, the wrong block at 33.

The term comparator now latches the first differing block with one
arithmetic statement per block (latch := latch + (latch = 0) * (block + 1)
* differs): no mask, no float, no width limit.

The array comparator compares each paired term whole in one statement
(its PRP bytes and left blocks, the deterministic part) and latches the
first unequal pair. It then calls compare_ore_block_256_term exactly once
on every path: on the latched pair, or, when every pair is equal or the
latched pair has a NULL element, on a NULL-free pair whose result is
discarded. The result is selected arithmetically. This removes the copy of
the block scan and hash from the array comparator, and every element is
read once per iteration. A malformed term paired with a NULL element now
raises too.

Measured on PostgreSQL 17.11 (aarch64), through the operator path:
single-term t = -0.78 and +0.48 (200k samples of 16), six-term t = +1.10
and -1.00 (50k samples of 8), and two new shapes at 100k samples of 8:
an equal prefix of different-length arrays against a first-term
difference, t = -0.45 and +0.79; a NULL after equal terms against a
first-term difference, t = +0.26 and +1.24. A six-term comparison now
takes about 18 µs either way (31 µs before this change); single-term
comparisons cost about 6% more than the original comparator.

Results are unchanged: 21,000 random arrays of 1-6 terms with NULL
elements and truncation, and random terms of 8 to 64 blocks, give exactly
what the original comparator gives, and the harness's 1,500 plaintext
checks pass.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

EQL ORE block comparator timing depends on where operands first differ

2 participants