Evidence, with its boundaries

What the tests establish.
And what they don’t.

Dated observations from public CI logs and benchmark documentation, reviewed on 1 October 2026. This is a snapshot, not a claim that every number tracks the latest release.

Passing tests support specific claims. They do not establish production capacity, external adoption, or long-term availability. Suites and matrix jobs can overlap, so there is no portfolio-wide test total.

CloudScale Backend

Public source snapshot: 26a7047de942

The decision: Use transaction-bound idempotency, balanced postings, and a transactional outbox. SQLite and PostgreSQL adapters share contracts, with integration testing against PostgreSQL.

The restore drill dumps and restores the database, rebuilds derived tables, and compares state. Its seeded dataset is deliberately small. A separate suite in the same job is not added to the primary count.

Recovery at seeded scale, not a sustained-availability claim.

FastAPI Microservices Platform

Public source snapshot: 8789684b1d29

The decision: Keep PostgreSQL as the only durable queue, without Redis or Kafka. Leased jobs use token-checked updates, with Standard Webhooks signatures and isolated delivery egress.

The million-job result exercises generated jobs and the PostgreSQL claim/finalize path on a laptop-class machine. It is not a million real HTTP deliveries or a production capacity result. A separate worker sample with a 5 ms mock receiver is documented in the benchmark. Scripted onboarding timings are not independent human usability evidence.

At-least-once delivery; the backlog-latency target was not met.

Code Quality Analyzer

Public source snapshot: 2dcaeb23dfa1

The decision: Analyze code offline and integrate into pull requests using SARIF, changed-line gates, and the companion cqa-action. Publish the analyzer as cqa-analyzer, not the unrelated code-quality-analyzer package on PyPI.

The selected observation is one deep-enabled Python job, not the sum of all OS and Python versions. Rules have documented limitations and fixtures. Passing those fixtures does not establish low false-positive rates on an arbitrary codebase.

Bounded static analysis, not proof of arbitrary program correctness.

AgentOS

Public source snapshot: 1e3bfb5884c2

The decision: Record durable workflow state in an event log, require approval for declared effects, and enforce an operator policy ceiling. An operator UI supports inspection and approval.

Core and UI results are separate scopes. Crash and network-failure CI jobs exercise recovery, but a committed workflow step and an external side effect are different boundaries. HMAC seals are tamper-evident under key protection, not tamper-proof. Only the public repository is referenced here.

External side effects depend on downstream idempotency or reconciliation.

Wallet Transfer Service

Public source snapshot: 360c87c9731b

The decision: Use transaction-bound idempotency keys and deterministic wallet lock ordering for a focused transfer service.

The suite runs against a real PostgreSQL service, including concurrent first-transfer and idempotent-retry scenarios. It establishes the invariants under test, not throughput or production use.

Transaction-design case study, not a production payments system.

Agent Skills

Public source snapshot: 0da9f0af0855

The decision: Package four tool-backed coding-agent workflows, including architecture snapshots and changed-line quality gates.

This entry describes the pinned public snapshot, not unpublished local changes. The selected job is one Python version of a four-version matrix. Script and contract tests establish tool behaviour; model-level validation needs a separate, scored evaluation. No marketplace adoption is claimed.

Script tests do not establish model-behaviour quality.