Python

Python — protocol and earlier observations

Measure complete Python requests, reusable-session overhead and binding-specific latency evidence.

This page records measurements of complete Python requests. They have their own protocol and are not part of the native C timing series.

For C microbenchmarks, use Native measurements by release. Session-memory reports remain available separately. Neither dataset is a Python benchmark.

Complete-request latency

Individual release measurements

Explore Python measurements for each release: choose the release, SMALL or LARGE, and the pass. Totals, input, solve, query and close timings are separate views. The September 27 campaign targets the 18 published releases from 0.1.0-alpha.3 through 0.12.0; only a completed campaign is imported.

Each immutable tag supplies its own Python binding and native engine. The V1 bindings have no reusable-session API: those cases are marked unavailable, never replaced by a modern binding or a simulated session. Releases containing Python Next use their own maelys_datalog_next; unified releases use their own maelys_datalog. The selected binding generation remains visible.

The same local Mac, Python interpreter and fixed CFFI tooling are used for both profiles, with native and extension compilation at -O2. All builds finish before measurement. Each available warm case keeps 501 samples after 50 warmup requests; the seven-symbol first-request cases use 31 fresh interpreters, excluding import and policy compilation. Four A/A passes precede forward/reverse rounds. Total and phase loops are different requests, not paired observations. Resource/GC telemetry accompanies the samples; no slow value is filtered or corrected.

The latest binary also runs a fixed three-request control. Its report states whether that coarse cost was detected in both rounds of every case/profile. This does not establish sensitivity to small effects. Old hosted observations below remain separate evidence: a new local run neither replaces their machine/protocol nor erases a previous tail alert.

A fast solver is only one part of an application request. Converting input, adding facts, querying the answer and releasing resources also take time.

ARCHITECTUREWhat a complete request measures
    • Build inputvalues + native conversions
    • Solveevaluate the complete EDB
    • Querypositive + negative decisions
    • Releaseresult and request resources

Reusable-session setup is measured separately. A convenience solve also creates and closes its private session, so those two paths must not be conflated.

Keep the two integration paths distinct:

  • Prepared session: load and prepare once, then evaluate successive requests. Session creation is outside the repeated-request measurement.
  • Convenience solve: the binding creates a private session for the call. Its setup and destruction are part of the complete request.

Native C, Python/CFFI and JavaScript/WASM must also be measured separately. A native C chart does not include interpreter objects, boundary conversions, browser work or network costs.

What the 0.12.0 Python evidence says

The release introduces deterministic native-call budgets and a repeated Python measurement protocol. A benchmark-only control runs three complete requests instead of one using the same baseline binary, to check that the protocol detects that deliberately larger cost. Detecting it does not prove sensitivity to every smaller regression.

The published candidate comparison measured f975be6 against 0.11.1. Most warm complete-request medians improved, but a SMALL-Release, seven-integer prepared-session case retained a p95 of 221.984 versus 89.587 microseconds in one round, and a much smaller difference in the other. The cause remains unresolved.

A later instrumented matrix records CPU time, scheduling/resource counters and GC intervals. It uses a different observation protocol and does not erase or reclassify the earlier result. Measurements of candidates are identified as such; they are not silently relabelled as fresh timings of every final release binary.

Python measurement protocol

  • Name the release and the exact measured commit. A pre-release candidate remains a candidate observation unless evidence covers the final runtime.
  • Record the machine, architecture, compiler/flags, profile, binding/Python versions, workload and whether setup is included.
  • Measure the same binary against itself first (A/A) to expose the noise floor, then alternate baseline and candidate runs (A/B).
  • Follow the declared protocol: the release's Python comparator uses minima below 10 microseconds and median/p95 otherwise. An A/B difference below its corresponding noise floor is indeterminate, not a win or proof of equality.
  • Show the tail as well as typical latency. p95 concerns the slower end of the sample distribution; it is not a worst-case bound.
  • Keep native, interpreter and instrumented/uninstrumented protocols separate. Do not directly compare absolute timings from different hosts or protocols.
  • Retain previous results, alerts and accepted uncertainties. A later quiet run does not retrospectively invalidate an earlier observation.

A measured improvement belongs to that workload and environment. It does not, by itself, establish the cause or guarantee the same gain in another application.

What you have learned

Working model

Measure the Python request, not only its solve
  • Include the binding cost

    Python inputs, native conversions, queries and cleanup contribute to a complete request.

  • Distinguish reusable and convenience paths

    A prepared-session request excludes one-time session setup; a convenience solve includes its private session’s setup and destruction.

  • Keep the protocol explicit

    Native-call budgets, A/A noise floors and timings answer different questions; instrumentation can change what is observed.

  • Retain the actual evidence

    Candidate comparisons and unresolved p95 observations must not be relabelled as guarantees for every final release.