Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Boruna

CI License: MIT Version Status: Stable

v3.0.0 is the current release. The 1.x line remains under long-term support — active through 2027-11-15, security through 2028-05-15. See docs/lts.md for support windows, deprecation policy, and security-backport SLAs.

Deterministic, policy-gated workflow execution for AI systems that must be auditable.


The problem

Most AI orchestration tools run workflows and return outputs. When something goes wrong — or when a regulator asks — there is no reliable way to answer: What exactly ran? What did the model see? What did it return? Can you prove it?

Boruna answers those questions by design.

Every Boruna workflow run produces a tamper-evident evidence bundle: a hash-chained audit log of every step executed, every capability invoked, every model response received. That bundle can be inspected, verified, and replayed — without network access, without a central server, without trusting anyone’s word.

This makes Boruna suited for teams building AI workflows that touch regulated data, make consequential decisions, or need a defensible audit trail.

What Boruna provides

  • DAG workflow execution — steps are .ax source files; the workflow is a workflow.json DAG definition (schema_version: 1 frozen at 1.0)
  • Capability enforcement — every side effect (LLM calls, HTTP, database, filesystem) is declared and policy-gated at the VM level
  • Evidence bundles — hash-chained tamper-evident logs, written automatically with --record. Optional AES-256-GCM envelope encryption for compliance-sensitive deployments. evidence inspect shows step output content for plaintext bundles.
  • Deterministic replay — re-execute any recorded workflow with identical outputs, verified by the VM
  • Approval gates — pause workflow execution for human review or external triggers before continuing
  • Diagnostics, auto-repair, and migration — boruna lang check, boruna lang repair, boruna migrate for .ax files and bundle/workflow upgrades
  • boruna new — interactive scaffold for new workflows from templates
  • 33 built-in functions — string (12), list (7), and map (7) operations plus type conversions and debug builtins (__builtin_string_*, __builtin_list_*, __builtin_map_*, …) available in every .ax file without imports
  • Import resolution — import "std-name" inlines libs/<name>/src/core.ax at compile time; 14 stdlib packages (the original 13 are 1.0-stable)
  • Four formal versioned specifications — .ax language 1.0, bytecode 1.0, evidence bundle format 1.0, workflow DAG schema 1.0 (all under docs/spec/)
  • MCP server — exposes 14 tools for AI coding agent integration (Claude Code, Cursor, Codex)

What Boruna is not

  • Not a general-purpose language or runtime (use Rust, Python, Go for that)
  • Not an LLM framework (use LangChain, LCEL, etc. for prompt engineering)
  • Not a cloud service (Boruna runs wherever you deploy it)
  • Not a key-management system (operators wire HSM / KMS integration themselves; bundle-encryption KEK lifecycle is operator-owned)

Install

Linux and macOS:

curl -fsSL https://raw.githubusercontent.com/escapeboy/boruna/master/install.sh | sh

Windows (PowerShell):

irm https://raw.githubusercontent.com/escapeboy/boruna/master/install.ps1 | iex

Both scripts pick the right build for your OS and CPU, check it against the release’s SHA256SUMS and refuse to install on a mismatch, then put boruna, boruna-mcp, boruna-pkg and boruna-orch in ~/.local/bin (Windows: %LOCALAPPDATA%\Programs\boruna\bin, added to your user PATH). Set BORUNA_VERSION=v3.4.0 to pin a version or BORUNA_INSTALL_DIR to choose the folder. Read install.sh / install.ps1 first if you prefer not to pipe a script into your shell.

Check the installation with boruna --version and boruna doctor.

Manual download with checksum check
# Linux x86_64 (musl — works on Alpine, Ubuntu, Debian, ...)
curl -fsSL https://github.com/escapeboy/boruna/releases/latest/download/SHA256SUMS -o SHA256SUMS
TARGET=x86_64-unknown-linux-musl
TAR=$(grep "$TARGET" SHA256SUMS | awk '{print $2}')
curl -fsSLO "https://github.com/escapeboy/boruna/releases/latest/download/$TAR"
grep "$TAR" SHA256SUMS | sha256sum -c -
tar -xzf "$TAR"
./boruna-*-${TARGET}/boruna --version

Windows (PowerShell):

$base = "https://github.com/escapeboy/boruna/releases/latest/download"
Invoke-WebRequest "$base/SHA256SUMS" -OutFile SHA256SUMS
$target = "x86_64-pc-windows-msvc"   # or aarch64-pc-windows-msvc on Windows on Arm
$zip = ((Select-String -Path SHA256SUMS -Pattern $target).Line -split "\s+")[1].TrimStart("*")
Invoke-WebRequest "$base/$zip" -OutFile $zip
# Compare this hash with the line for $zip in SHA256SUMS:
(Get-FileHash $zip -Algorithm SHA256).Hash.ToLower()
Expand-Archive $zip -DestinationPath .
.\boruna-*-$target\boruna.exe --version

Supported platforms

Every platform below is built and its full test suite is run natively on a real machine of that type in CI (no emulation), together with an example workflow and evidence verify.

PlatformRelease assetTested natively in CI
Linux x86_64 (musl)x86_64-unknown-linux-musl .tar.gzyes (tests run on a glibc build; the musl release binary is cross-built)
Linux arm64 (musl)aarch64-unknown-linux-musl .tar.gzyes (glibc build; the musl release binary is cross-built)
macOS Apple Siliconaarch64-apple-darwin .tar.gzyes
macOS Intelx86_64-apple-darwin .tar.gzyes
Windows x64x86_64-pc-windows-msvc .zipyes
Windows on Armaarch64-pc-windows-msvc .zipyes

Other targets (FreeBSD, 32-bit, RISC-V and so on) are not built or tested; build from source and run cargo test --workspace to see whether they work for you. See docs/releasing.md.

Or build from source:

git clone https://github.com/escapeboy/boruna
cd boruna
cargo build --workspace --release

Quickstart

# Run a workflow
boruna workflow run examples/workflows/llm_code_review --policy allow-all --record

# Verify the evidence bundle
boruna evidence verify .boruna/runs/<run-id>/

→ Full Quickstart — 10 minutes, ends with a verified evidence bundle.

Example workflows

WorkflowPatternWhat it shows
LLM Code ReviewLinear, 3 stepsLLM capability, data flow, evidence recording
Document ProcessingFan-out, 5 stepsParallel steps, multi-input merge
Customer Support TriageApproval gateHuman-in-the-loop, conditional pause, audit trail

Each example runs in demo mode (no external services) and produces a verifiable evidence bundle.

How the evidence guarantee works

workflow.json  →  DAG Validator  →  Step Runner
                                        ↓
                                   .ax source
                                        ↓
                                   Compiler → Bytecode
                                        ↓
                                   VM (capability gateway)
                                        ↓
                                   EventLog entry (CapCall + CapResult)
                                        ↓
                              Hash-chained audit log

boruna evidence verify <bundle>
  → Chain integrity: VALID
  → All step hashes: MATCH
  → Verification: PASSED

Every CapCall (including LLM calls) is logged with its full response. The log is SHA-256 hash-chained from a genesis entry containing the workflow definition hash. Modification of any entry breaks the chain.

Architecture

Boruna is a Rust workspace with 10 production crates plus a benches/ member:

CratePurpose
boruna-orchestratorWorkflow engine, DAG execution, evidence bundles
boruna-vmBytecode VM, capability gateway, actor system, replay
boruna-compilerLexer, parser, type checker, code generator
boruna-bytecodeOpcodes, Module, Value, Capability definitions
boruna-frameworkElm-architecture runtime, test harness
boruna-effectLLM integration, prompt management, caching
boruna-cliCLI binary (boruna)
boruna-toolingDiagnostics, repair, trace-to-tests, templates
boruna-pkgPackage registry, resolver, lockfiles

1175+ tests across 11 workspace members. cargo test --workspace — all pass.

Documentation

QuickstartBuild, run a workflow, inspect evidence
Concepts: DeterminismWhy and how determinism is enforced
Concepts: CapabilitiesSide effect declaration and policy gating
Concepts: Evidence BundlesHash-chained audit logs and replay
Guide: First WorkflowBuild a workflow from scratch
Guide: MigrationUpgrade legacy bundles and workflow files
Spec: .ax Language 1.0Formal language specification
Spec: Workflow DAG 1.0workflow.json schema
Spec: Evidence Bundle 1.0Bundle format + encryption envelope
Reference: CLIAll boruna commands
LTS contractSupport windows + deprecation policy for 1.x
PerformanceBaseline numbers + 1.x performance budget
StabilityWhat is stable, experimental, and planned
Roadmap0.2.0 → 1.0.0 → 1.x
LimitationsReal constraints, stated honestly
FAQCommon questions
All docs →Full documentation index

Status

Boruna is at v3.0.0 — the release that removes the entire HTTP / serving / distributed-execution layer. Gone are the distributed coordinator, distributed workers, active-active HA and coordinator mTLS, the three web UIs (workflow dashboard, evidence web viewer, approval console), and the serve cargo feature and its server dependencies. What remains is a local deterministic engine and CLI: compiler → capability-gated VM → orchestrator (runner, persistence, audit) → tamper-evident evidence bundles. Approval and external-trigger gates are still handled locally via boruna workflow approve/reject/trigger plus resume. This is a breaking release — the coordinator, dashboard, worker, and evidence serve CLI commands, the --coordinator / --coord-token flags, and the serve feature are removed — so review the 3.0.0 entry in CHANGELOG.md. The core execution engine, evidence bundles, and four formal versioned specifications (.ax language, bytecode, workflow DAG, evidence bundle) remain feature-complete; the 1.x LTS line continues per docs/lts.md.

The project is suited for evaluation, internal tooling, and audit-sensitive AI pipelines. Operator action: validate the docs/PERFORMANCE.md budget against your workload, and review docs/limitations.md for known constraints. External security audit booking is the Q4 2026 commitment in lts.md.

See docs/stability.md for the stability tier breakdown.

For coding agents

Boruna exposes an MCP server for AI coding agent integration. See AGENTS.md for integration instructions and the tool reference.

Contributing

See CONTRIBUTING.md. The short version: open an issue, implement with tests, run cargo test --workspace + cargo clippy + cargo fmt, add a CHANGELOG entry, open a PR.

License

MIT — Copyright 2026 Boruna Contributors

Quickstart

Get from zero to a running workflow with an evidence bundle in about 10 minutes.

Prerequisites

  • Rust 1.75+ (rustup update stable)
  • Git

1. Build

git clone https://github.com/escapeboy/boruna
cd boruna
cargo build --workspace

This builds all 11 workspace members (10 production crates + benches/). Expect 1-2 minutes on first build.

No Rust? Install the prebuilt binaries instead (see Install), still clone the repository for the examples, and write boruna wherever this guide says cargo run --bin boruna --.

2. Run hello world

cargo run --bin boruna -- run examples/hello.ax

Expected output:

Hello, Boruna!

This compiles hello.ax to bytecode and runs it on the VM. No capabilities are needed.

3. Run a workflow

Boruna workflows are DAGs — directed acyclic graphs of steps. Each step is a .ax file that compiles independently and runs in isolation.

Run the LLM code review workflow (in demo mode — no real LLM calls):

cargo run --bin boruna -- workflow run examples/workflows/llm_code_review \
  --policy allow-all

Expected output:

Running workflow: llm_code_review

  [1/3] fetch_diff    → ok
  [2/3] analyze       → ok
  [3/3] report        → ok

Workflow completed in 0.03s

Three steps ran in topological order: fetch a diff, analyze it, produce a report. In demo mode the steps return representative data without calling any external services.

4. Record an evidence bundle

Add --record to capture a tamper-evident log of the run:

cargo run --bin boruna -- workflow run examples/workflows/llm_code_review \
  --policy allow-all --record

Expected output:

Running workflow: llm_code_review

  [1/3] fetch_diff    → ok
  [2/3] analyze       → ok
  [3/3] report        → ok

Workflow completed in 0.03s
Bundle written to: .boruna/runs/20260319-120000-abc12/

5. Inspect the evidence bundle

cargo run --bin boruna -- evidence inspect .boruna/runs/20260319-120000-abc12/

Expected output:

Run ID:     20260319-120000-abc12
Workflow:   llm_code_review
Started:    2026-03-19T12:00:00Z
Completed:  2026-03-19T12:00:00Z
Policy:     allow-all
Steps:      3 completed, 0 failed

Step Results:
  fetch_diff   → ok  (0.0s)
  analyze      → ok  (0.0s)
  report       → ok  (0.0s)

Chain:      valid (3 entries, no gaps)

6. Verify the evidence bundle

cargo run --bin boruna -- evidence verify .boruna/runs/20260319-120000-abc12/

Expected output:

Chain integrity: VALID
All step hashes: MATCH
Environment fingerprint: PRESENT
Verification: PASSED

The hash chain is unbroken. No step output was modified. This is what makes Boruna useful for audit: you can present this bundle, and anyone with the boruna binary can verify it independently.

What you just saw

  • Deterministic execution: same workflow definition → same outputs, every time
  • Capability policy: --policy allow-all controls what side effects are permitted
  • Evidence bundle: a tamper-evident directory written alongside every recorded run
  • Independent verification: evidence verify needs no network access, no central server

Try the other workflows

# Document processing with fan-out parallelism
cargo run --bin boruna -- workflow run examples/workflows/document_processing \
  --policy allow-all --record

# Customer support triage with an approval gate
cargo run --bin boruna -- workflow run examples/workflows/customer_support_triage \
  --policy allow-all --record

Run the tests

cargo test --workspace --features boruna-cli/serve

1175+ tests across all workspace members. All should pass.

Next steps

Your First Workflow

This guide walks through creating a simple two-step workflow from scratch, running it, and inspecting the evidence it produces.

Time required: ~15 minutes Prerequisites: Boruna built (cargo build --workspace)

What you’ll build

A workflow that fetches a number and doubles it — two steps connected in sequence. Simple enough to fit in your head, realistic enough to demonstrate every core concept: DAG definition, .ax step code, capability policies, and evidence.

Step 1: Create the workflow directory

mkdir -p my-first-workflow/steps
cd my-first-workflow

Step 2: Define the workflow DAG

Create workflow.json:

{
  "id": "double-it",
  "name": "Double It",
  "description": "Fetch a number, double it, return the result.",
  "steps": [
    {
      "id": "fetch",
      "name": "Fetch Number",
      "source": "steps/fetch.ax",
      "capabilities": []
    },
    {
      "id": "double",
      "name": "Double It",
      "source": "steps/double.ax",
      "capabilities": [],
      "depends_on": ["fetch"]
    }
  ],
  "edges": [
    { "from": "fetch", "to": "double" }
  ]
}

The depends_on field defines the DAG edge. double runs after fetch completes.

Step 3: Write the step files

steps/fetch.ax — returns a number (demo mode; live mode would call net.fetch):

fn get_number() -> Int {
    42
}

fn main() -> Int {
    let n: Int = get_number()
    n
}

steps/double.ax — doubles the input:

fn double(n: Int) -> Int {
    n * 2
}

fn main() -> Int {
    let input: Int = 42
    let result: Int = double(input)
    result
}

Step 4: Validate the workflow

Before running, validate the DAG structure:

cd ..
cargo run --bin boruna -- workflow validate my-first-workflow/

# Expected output:
# Workflow 'double-it': valid
# Steps: 2
# Edges: 1
# Topological order: fetch → double
# No cycles detected.

Validation checks: all referenced step files exist, the graph is acyclic, and every depends_on refers to a real step.

Step 5: Run the workflow

cargo run --bin boruna -- workflow run my-first-workflow/ --policy allow-all

# Expected output:
# Running workflow: double-it
#
#   [1/2] fetch     → ok
#   [2/2] double    → ok
#
# Workflow completed in 0.01s
# Output: 84

Step 6: Run with evidence recording

cargo run --bin boruna -- workflow run my-first-workflow/ --policy allow-all --record

# Expected output:
# Running workflow: double-it
# ...
# Workflow completed in 0.01s
# Bundle written to: .boruna/runs/20260319-120000-abc12/

Step 7: Inspect the evidence

cargo run --bin boruna -- evidence inspect .boruna/runs/20260319-120000-abc12/

# Expected output:
# Run ID:     20260319-120000-abc12
# Workflow:   double-it
# Started:    2026-03-19T12:00:00Z
# Completed:  2026-03-19T12:00:00Z
# Policy:     allow-all
# Steps:      2 completed, 0 failed
#
# Step Results:
#   fetch    → ok  (0.0s)
#   double   → ok  (0.0s)
#
# Chain:      valid (2 entries, no gaps)

Step 8: Verify the evidence bundle

cargo run --bin boruna -- evidence verify .boruna/runs/20260319-120000-abc12/

# Expected output:
# Chain integrity: VALID
# All step hashes: MATCH
# Environment fingerprint: PRESENT
# Verification: PASSED

That’s a complete workflow: defined, validated, executed, recorded, and verified.

What just happened

  1. The workflow runner read workflow.json and sorted steps topologically.
  2. Each .ax file was compiled to bytecode and run on the VM.
  3. Every step’s output was captured and written to the evidence bundle.
  4. The audit log was hash-chained so no entry can be modified undetected.
  5. evidence verify confirmed the chain was unbroken.

Next steps

Limitations

Boruna has real constraints. This document describes them clearly, so you can make an informed decision about whether it fits your use case.

Language limitations

Values are immutable; only let mut bindings can be rebound. Records, lists and maps are never changed in place; state transitions use record spread (State { ..old, field: new_value }). A let mut binding can point to a new value (total = total + 1), which is enough for counters and accumulators but not for shared mutable state.

Loops are basic. There are while and for x in list loops, but no break, continue or loop, and for only iterates lists. The step limit stops a loop that never ends, and deep recursion can also hit it.

Some type checks are not enforced yet. Assigning a value of a different type to a let mut binding, using a non-Bool while condition, and reassigning a binding declared without mut all compile today (the last one is warning E010). See the language spec §4.5.

No generics. Types in .ax are concrete at definition time. There is no generic type system. This keeps the language simple but limits abstraction.

No imports. .ax files cannot import other .ax files. Standard library access is through the package system, not file imports. Large workflows are composed at the workflow DAG level, not at the language level.

String processing is limited. The standard library provides basic string operations, but .ax is not designed for complex text transformation. Use LLM capabilities for natural language processing; use compiled Rust for heavy text manipulation.

VM limitations

Step limit is a blunt instrument. The --step-limit flag prevents runaway execution but does not provide fine-grained CPU time control. A step that does 10M arithmetic operations may run longer than expected before hitting the limit.

Wave-based concurrent execution, not full DAG parallelism. The runner processes steps in topological waves with --concurrency <N> workers per wave (sprint 0.3-S4). A slow step at level N blocks fast steps at level N+1 even if they don’t actually depend on the slow one. A full DAG-based scheduler (no wave boundaries) is not yet implemented.

Actor system is single-process. Actors run in the same OS process with round-robin scheduling. There is no distributed actor execution.

Memory is unbounded within a step. The VM does not enforce memory limits on individual step execution. A step that allocates large lists or maps may use significant memory.

Capability limitations

The HTTP handler requires a feature flag. Real HTTP calls require building with --features boruna-cli/http. This is a build-time decision, not a runtime one.

LLM calls use Bring Your Own Handler (BYOH). The llm.call capability is declared and enforced by the VM, and boruna-effect provides prompt / cache / context primitives — but no default network-calling handler ships in core. Wire your provider (OpenAI, Anthropic, vLLM, Ollama, custom router) by implementing CapabilityHandler and passing it to CapabilityGateway::with_handler. See LLM Integration Guide for the contract, examples, and rationale. A reference OpenAI handler lives at examples/llm_handlers/openai/.

No streaming. Capability calls are synchronous and blocking. LLM responses must complete before the step continues. This is unsuitable for streaming chat interfaces.

SSRF protection is allowlist-based. The HTTP handler rejects private IP ranges and localhost, but allowlisting specific domains requires extending the NetPolicy struct. There is no UI for this.

Workflow limitations

Wall-clock-keyed enforcement is non-deterministic on failure. Limits like max_wall_ms are wall-clock-keyed: a workflow that completes within budget produces deterministic output, but one that times out may finish on a fast machine and time out on a slow one. Documented per integrator surface (limits, OTel spans).

Evidence and audit limitations

Evidence bundles are local files; remote storage is operator-owned. Evidence bundles write to <data-dir>/runs/<run-id>/. Pluggable storage adapters (S3 / object storage / document store) are roadmap 0.7.x or 1.x. Today, ship bundles to remote storage with your own pipeline (rsync, S3 upload, etc.).

LLM response reproducibility is not guaranteed. Evidence bundles capture LLM responses for replay, but if the LLM provider changes their model weights, a replay may produce different outputs if the real capability is used. Replay with recorded responses (sprint 0.5-S7 of FleetQ track) is always reproducible.

Evidence bundle encryption KEK is operator-managed. Sprint W6-B added envelope AES-256-GCM encryption — operators supply the KEK via env var or CLI flag. Boruna does not ship key management (no HSM/KMS integration); KEK lifecycle (storage, rotation, sealing) is the operator’s responsibility. Key rotation tooling is roadmap.

Plaintext bundle.json metadata. Even with --encrypt-bundle, the top-level bundle.json manifest is plaintext (chicken-and-egg with the wrapped DEK). It carries format_version, boruna_version, run_id, workflow_hash, and the wrapped DEK envelope. Run identifiers may be visible to a bundle inspector even when payload bytes are encrypted.

Operational limitations

Multi-tenancy is environment-namespaced, not cryptographically isolated. The --env flag (sprint 0.4-S14) namespaces the data-dir and Prometheus labels per-environment. This separates run histories but does not provide cryptographic isolation between tenants — that requires OS-level separation or per-tenant deployments.

Minimum Rust version: 1.75.0. Teams running older Rust toolchains will need to upgrade.

What to do if these limitations block you

File an issue at https://github.com/escapeboy/boruna. Limitations that frequently block real use cases will be prioritized in the roadmap.

See also: Roadmap, Stability

Frequently Asked Questions

What problem does Boruna solve?

Most AI workflow orchestration tools are wrappers around LLM API calls with some retry logic and prompt templates. They work fine for demos. They break down when you need to answer questions like: “What exactly did the model see?”, “Can I prove this workflow ran correctly?”, “If I re-run this tomorrow, will I get the same result?”

Boruna is built for teams that need answers to those questions — because they operate in regulated environments, because their AI workflows make consequential decisions, or because they’ve been burned by non-reproducible outputs.

The core value proposition: every workflow run produces a tamper-evident evidence bundle. You can verify it independently, replay it exactly, and present it in an audit.

How is Boruna different from LangChain, LlamaIndex, or similar frameworks?

LangChain and similar frameworks are excellent for building LLM-powered applications quickly. They are not designed for auditability or deterministic replay.

Boruna’s design priorities are different:

LangChainBoruna
Primary goalLLM integrationDeterministic execution
Side effectsImplicitDeclared and gated
Audit trailNot built-inHash-chained, always
ReplayNot supportedFirst-class
Policy enforcementNoneCapability-based
LanguagePython.ax (custom, Rust VM)

These are different tools for different requirements. A team building a chatbot should probably use LangChain. A team running AI-assisted financial analysis in a regulated environment should look at Boruna.

How is Boruna different from Temporal or similar workflow engines?

Temporal is a durable workflow engine with excellent reliability guarantees. It is not specifically designed for AI workflows or LLM governance.

Boruna focuses on:

  • Capability policy: declaring and enforcing what LLM calls and network requests are permitted
  • Evidence bundles: cryptographic proof of what ran and what the model returned
  • Determinism as a contract: same inputs → same outputs, enforced by the VM
  • The .ax language: pure, typed, auditable step code that compiles to bytecode

Temporal is a better choice if you need multi-step business processes with human tasks, timers, and at-least-once execution guarantees. Boruna is a better choice if governance and auditability of AI-specific workflows is the primary concern.

Is .ax Turing-complete?

Yes. .ax supports recursion and is Turing-complete in the theoretical sense. In practice, the --step-limit flag on boruna run enforces a bound on VM steps, which prevents runaway execution in production workflows.

Can I call external APIs and LLMs?

Yes, but they must be declared as capabilities. A step that calls an LLM must annotate its function with !{llm.call}. A step that makes HTTP requests must declare !{net.fetch}.

In demo mode (default), capability calls are stubbed. In --live mode (with the http feature), real HTTP requests are made against the SSRF-protected handler.

Does Boruna guarantee that LLM outputs are reproducible?

No. LLMs are probabilistic. Running the same prompt twice will produce different outputs. Boruna cannot change this.

What Boruna guarantees: the LLM response is recorded in the evidence bundle. If you replay the workflow from its evidence bundle, the recorded response is substituted — so the replay is deterministic even if the original call wasn’t.

This lets you verify that a workflow produced a specific output on a specific run, without requiring the LLM to reproduce the response.

What is the evidence bundle format?

A directory containing:

  • manifest.json — run metadata
  • audit_log.json — hash-chained step execution log
  • events/event_log.json — full capability call stream
  • steps/<step-id>.input and .output — step I/O
  • env_fingerprint.json — runtime environment snapshot

See Evidence Bundles for the full specification.

Can I use Boruna without the .ax language?

Not currently. The VM, capability enforcement, and determinism guarantees are tied to .ax bytecode. Compiling other languages to Boruna bytecode is technically possible but not a current project goal.

What Rust version does Boruna require?

Minimum supported Rust version: 1.75.0 (stable). No nightly features are required.

Is there a hosted version?

Not yet. Boruna runs locally or in whatever environment you deploy it to. A hosted platform is on the long-term roadmap.

How do I report a security vulnerability?

See SECURITY.md. Use GitHub Security Advisories for responsible disclosure. Do not file a public issue.

Is Boruna production-ready?

Boruna is at 1.3.0, on the 1.x LTS line. The core engine, workflow DAG, evidence bundles, capability enforcement, and all 13 stdlib packages are stable and LTS-protected. External security audit is booked for Q4 2026. See Stability for the full maturity assessment and lts.md for the support contract.

How do I contribute?

See CONTRIBUTING.md. The short version: open an issue, fork, implement, run cargo test --workspace + cargo clippy + cargo fmt, and open a PR with a CHANGELOG entry.

Determinism in Boruna

Determinism is Boruna’s foundational guarantee: given the same workflow definition, the same inputs, and the same capability responses, execution produces identical outputs every time, on every machine.

This is not a convention or a best-effort property. It is structurally enforced by the runtime.

Why determinism matters

AI systems are inherently probabilistic. LLMs return different outputs on repeated calls. Network responses vary. Clocks differ between machines. In ad hoc orchestration this is accepted as normal. In regulated environments, it creates a compliance problem: you cannot audit what you cannot reproduce.

Boruna’s answer is to treat non-determinism as a capability, not a default. Every source of external state must be explicitly declared, gated by policy, and logged in the evidence bundle. The pure .ax computation in between is deterministic by construction.

This makes it possible to:

  • Reproduce any workflow execution exactly from its evidence bundle
  • Prove that a recorded run matches a given workflow definition
  • Detect if model behavior has drifted between runs
  • Satisfy audit requirements without manual log correlation

The determinism boundary

Boruna draws a hard boundary between pure computation and effects.

Pure (deterministic):

  • All .ax expressions and functions
  • Record and enum operations
  • Pattern matching, conditionals, loops
  • Function calls (including recursive)
  • Actor message passing (scheduling order is recorded)

Effects (non-deterministic, capability-gated):

  • Network requests (net.fetch)
  • LLM calls (llm.call)
  • Time reads (time.now)
  • File system access (fs.read, fs.write)
  • Database queries (db.query)
  • Random number generation (random)

Effects are declared on functions using capability annotations:

fn fetch_data(url: String) -> String !{net.fetch} {
    // implementation (live mode)
}

Without the !{net.fetch} annotation, a function cannot perform network I/O — the VM enforces this at the bytecode level.

How the EventLog works

Every capability call is intercepted by the CapabilityGateway and written to the EventLog as a CapCall/CapResult pair:

CapCall  { capability: "net.fetch", args: ["https://api.example.com/data"] }
CapResult{ capability: "net.fetch", value: "{\"status\": 200, ...}" }

The EventLog also captures actor lifecycle events (ActorSpawn, MessageSend, MessageReceive, SchedulerTick) so multi-actor scheduling is fully reproducible.

When --record is passed to workflow run, the EventLog is written into the evidence bundle.

Replay verification

Replay works by substituting recorded CapResult values instead of making real calls:

  1. Record: Run the workflow with live capabilities. Save the EventLog.
  2. Replay: Run the same bytecode. When a CapCall is encountered, return the recorded result instead of executing the effect.
  3. Verify: The CapCall sequence must match exactly — same capability, same arguments, same order.

If verification fails, either the workflow is non-deterministic (a bug) or an external value leaked into the pure core.

# Run and record
boruna workflow run examples/workflows/llm_code_review --policy allow-all --record

# Verify the evidence bundle
boruna evidence verify .boruna/runs/<run-id>/

BTreeMap, not HashMap

One concrete implication of the determinism guarantee: all ordered iteration in Boruna’s Rust implementation uses BTreeMap (sorted by key) rather than HashMap (random iteration order). This ensures that serialized outputs and logged values are identical across runs and platforms.

When writing .ax code, Map literals also use deterministic ordering.

Determinism boundaries and guarantees

GuaranteeScope
Same bytecode + same EventLog → identical outputVM, framework
Actor scheduling order is reproducibleVM actor system
Package resolution is locked (SHA-256 content hashes)boruna-pkg
Workflow step execution order follows DAG topologyboruna-orchestrator
Evidence bundle integrity (hash-chained log)audit module

Not guaranteed:

  • LLM output reproducibility (LLMs are probabilistic; Boruna records responses but cannot force models to repeat them)
  • Wall-clock timing of individual steps
  • External service behavior between record and replay

See also: Limitations, Evidence Bundles

Capabilities

A capability is an explicit permission for a workflow step to perform a side effect. No capability, no side effect — the VM enforces this unconditionally.

The eleven capabilities

The 1.0 capability set is frozen in crates/llmbc/src/capability.rs::Capability::ALL. A capability_set_hash, derived from the (name, version) tuples of this set, gives the capability surface a stable identity — reported by the boruna_capability_list MCP tool for compatibility checks.

CapabilityEffectExample use
net.fetchHTTP requestsCalling external APIs, webhooks
llm.callLLM inferenceGPT-4, Claude, local models (BYOH — see guides/llm-integration.md)
time.nowCurrent timestampTimestamping records
randomRandom numbersSampling, tie-breaking
fs.readFile system readsLoading documents, configs
fs.writeFile system writesWriting reports, outputs
db.queryDatabase accessReading/writing records
ui.renderUI surface renderingFramework view output (Elm-architecture apps)
actor.spawnActor creationSpawning parallel agents
actor.sendInter-actor messagingCoordinating actor state
step.inputRead a workflow step’s resolved inputsPer-step step_input("upstream_step") builtin

Declaring capabilities

Capabilities are declared on functions using the !{...} annotation:

fn call_model(prompt: String) -> String !{llm.call} {
    // live mode: calls LLM
}

fn fetch_and_parse(url: String) -> String !{net.fetch} {
    // live mode: makes HTTP request
}

A function without a capability annotation is pure — it cannot perform I/O and its output depends only on its inputs.

Policies

A policy is a set of allowed capabilities. It is specified at runtime, not in the workflow definition. This separation means the same workflow can run in restricted mode during testing and with full capabilities in production.

Built-in policies:

PolicyAllowed capabilities
allow-allAll 11 capabilities
deny-allNone
defaultNone (same as deny-all)

Pass a policy on the CLI:

boruna workflow run my-workflow/ --policy allow-all

In demo mode (no --live flag), capability calls are stubbed or skipped. In live mode, the policy is enforced against every capability call.

Capability enforcement in the VM

Every capability call in compiled bytecode goes through the CapabilityGateway:

  1. The VM encounters a capability opcode.
  2. The gateway checks the active policy.
  3. If the capability is not allowed, the VM returns a CapabilityDenied error immediately.
  4. If allowed, the call is dispatched to the registered handler (real HTTP, LLM client, etc.).
  5. The call and its result are written to the EventLog.

This enforcement happens at the bytecode level, before any handler executes. There is no way to bypass it from .ax code.

Capabilities in workflow definitions

Workflow steps declare their required capabilities in workflow.json. This makes the capability surface area visible before execution:

{
  "steps": [
    {
      "id": "analyze",
      "source": "steps/analyze.ax",
      "capabilities": ["llm.call"]
    }
  ]
}

The workflow validator checks that declared capabilities are consistent with the policy before the workflow runs.

Capabilities in package manifests

Standard libraries that require capabilities declare them in package.ax.json:

{
  "name": "std-http",
  "capabilities_required": ["net.fetch"]
}

This gives visibility into the transitive capability requirements of a dependency graph.

See also: Determinism, Policies

Evidence Bundles

An evidence bundle is the tamper-evident record of a workflow execution. It contains everything needed to prove what ran, when it ran, what inputs were used, what capabilities were invoked, and what outputs were produced.

Evidence bundles are the mechanism through which Boruna supports compliance, audit, and replay.

What an evidence bundle contains

.boruna/runs/<run-id>/
├── manifest.json          # Run metadata: workflow ID, start time, policy, step list
├── audit_log.json         # Hash-chained log of every step execution
├── events/
│   └── event_log.json     # Full CapCall/CapResult/actor event stream
├── steps/
│   ├── <step-id>.input    # Step input values (serialized)
│   └── <step-id>.output   # Step output values (serialized)
└── env_fingerprint.json   # Runtime environment: OS, Boruna version, hash of workflow def

Hash-chained audit log

The audit log is hash-chained: each entry includes the SHA-256 hash of the previous entry. This makes it impossible to insert, delete, or modify a log entry without breaking the chain.

Each log entry records:

  • Step ID and source file hash
  • Start time and end time
  • Policy in effect
  • Capability calls made
  • Output value hash
  • Previous entry hash

The chain starts with a genesis entry that includes the workflow definition hash and the environment fingerprint.

Generating a bundle

Pass --record to workflow run:

boruna workflow run examples/workflows/llm_code_review \
  --policy allow-all \
  --record

# Output:
# Bundle written to: .boruna/runs/20260315-143022-abc4d/

Inspecting a bundle

# Summary view
boruna evidence inspect .boruna/runs/20260315-143022-abc4d/

# Full JSON output
boruna evidence inspect .boruna/runs/20260315-143022-abc4d/ --json

Example output:

Run ID:     20260315-143022-abc4d
Workflow:   llm_code_review
Started:    2026-03-15T14:30:22Z
Completed:  2026-03-15T14:30:31Z
Policy:     allow-all
Steps:      3 completed, 0 failed

Step Results:
  fetch_diff   → ok  (0.1s)
  analyze      → ok  (6.8s)  [llm.call: 1 invocation, 312 tokens]
  report       → ok  (0.0s)

Chain:      valid (3 entries, no gaps)

Verifying a bundle

boruna evidence verify .boruna/runs/20260315-143022-abc4d/

# Output:
# Chain integrity: VALID
# All step hashes: MATCH
# Environment fingerprint: PRESENT
# Verification: PASSED

Verification checks:

  1. Hash chain is unbroken from genesis to final entry
  2. Step output hashes match recorded values
  3. Workflow definition hash matches the definition on disk
  4. Environment fingerprint is present and well-formed

A failed verification means either the bundle was tampered with, or the workflow definition changed since the run.

Replaying from a bundle

The evidence bundle contains everything needed to re-execute the workflow with the same capability responses:

boruna workflow run examples/workflows/llm_code_review \
  --replay .boruna/runs/20260315-143022-abc4d/ \
  --verify

In replay mode, LLM calls, HTTP requests, and all other effects return their recorded responses instead of hitting real services. The --verify flag checks that the replayed execution produces the same output hashes.

Compliance relevance

Evidence bundles address several audit requirements directly:

  • What ran: Workflow definition hash is recorded
  • What model was called: LLM capability calls logged with model identifier
  • What the model returned: CapResult entries preserve full responses
  • Who approved it: Approval gate decisions are logged as step transitions
  • Was it tampered with: Hash chain verification detects any modification

For regulated workflows (financial, healthcare, legal), evidence bundles can be exported and stored in a document management system alongside the artifacts they describe.

See also: Replay, Compliance

Evidence Bundle Threat Model

This document states, honestly and without marketing, what an evidence bundle proves and — just as important — what it does not prove. It is written in the SLSA spirit: enumerate the threats, name the concrete mitigation, and name the residual gap that remains after the mitigation.

If you take one sentence away, take this:

An evidence bundle lets you prove that the record was not altered after it was sealed, and that the record is internally consistent under replay. It does not, and cannot, prove that the record was true when it was written.

Everything below is an elaboration of that single distinction.


1. Three properties people conflate

Compliance conversations routinely blur three different guarantees. Boruna delivers the first, delivers the second only under stated conditions, and deliberately does not claim the third.

PropertyPlain-English meaningDoes Boruna provide it?
Tamper-evidenceIf someone changes the sealed record, a verifier can detect the change.Yes — this is the core guarantee.
Tamper-proofingChanging the sealed record is impossible.No. Nothing on a general-purpose filesystem is tamper-proof; a holder of the bytes can always rewrite them. Boruna makes tampering detectable, not impossible.
Non-repudiationThe party who produced the record cannot later deny producing it.Partial, and only with a signed bundle under a pinned key (see §4). Even then it attests who sealed the bytes, not whether the bytes are true.

Keep these separate. Most overclaiming comes from quietly upgrading “tamper-evident” to “tamper-proof”, or from treating a signature as proof of truth rather than proof of origin.


2. What a bundle actually contains

A sealed bundle (orchestrator/src/audit/evidence.rs, BundleManifest) carries, at minimum:

  • run_id, workflow_name, workflow_hash, policy_hash
  • audit_log_hash — the head of a hash-chained event log
  • file_checksums — SHA-256 of every component file (workflow, policy, audit log, env fingerprint, per-step outputs, and optional intents.json / model_invoking_steps.json)
  • env_fingerprint — OS / arch / Boruna version, self-reported
  • bundle_hash — SHA-256 over the manifest itself (excluding bundle_hash and signature), which therefore commits to every file_checksums entry and the audit_log_hash
  • optional encryption — AES-256-GCM envelope metadata
  • optional signature — an ed25519 signature over bundle_hash

The integrity contract enforced by verify_bundle (orchestrator/src/audit/verify.rs): every file’s on-disk SHA-256 must match file_checksums; the audit-log chain must be unbroken (entry_hash = SHA-256(prev_hash || event_json)); the chain head must equal audit_log_hash; and all required components must be present. For the full on-disk contract see Evidence Bundles and the format spec at docs/spec/evidence-bundle-1.0.md.


3. Threats and mitigations

ThreatWhat Boruna doesResidual gap
A third party edits a bundle file after it was sealed.The hash chain plus file_checksums plus bundle_hash make a naive edit detectable: any changed byte fails its SHA-256 check, and any spliced/removed audit entry breaks the chain. verify_bundle reports the failing check.A naive edit is caught by plain verify. A motivated attacker who holds the whole bundle can rewrite the file and recompute every checksum and recompute bundle_hash so the bundle is internally self-consistent — this defeats plain verify (documented in-code as the “F1 weakness”). Closing it requires an external anchor or a signature under a pinned key — see §4. Tamper-evidence, not tamper-proofing.
The recorder/producer is malicious and seals a FALSE record at write time.Nothing. The bundle faithfully seals whatever the producer fed it. A signature (if present) attests which key sealed these bytes — the producer’s identity — not that the sealed facts are true.Not prevented, by construction. Garbage-in is sealed as faithfully as truth-in. Evidence bundles are a tamper-evidence mechanism, not a truth oracle. Detecting a lying producer requires controls outside the bundle (independent corroboration, dual control over the recorder, a trusted execution environment — see §5).
The signing key is compromised.With a valid key an attacker can forge or backdate the entire bundle, sign it, and it will verify under that key — hash-chaining is single-writer and provides no defense once the writer’s key is held. Pinning a trusted_pubkey at verify time limits acceptance to a specific key, so a different attacker key is rejected.If the legitimate key itself is stolen, pinning does not help — the forged bundle carries the pinned key. There is no revocation, no key rotation history, and no witnessed record of when a signature was made. Mitigation direction (not yet implemented): anchoring signatures in an append-only transparency log and/or keyless, identity-bound signing, so a signature is bound to a witnessed moment and a verifiable identity rather than to a long-lived secret.
Backdating — sealing a record now but claiming it was produced earlier.The manifest carries started_at / completed_at / created_at timestamps, but these are self-reported wall-clock values written by the producer. Nothing external witnesses them.No trusted timestamp today. A producer (or a key holder) can set these fields to any value. Mitigation direction (not yet implemented): anchoring the bundle_hash in an external append-only log (e.g. a Rekor-style transparency log) at seal time, so the earliest-existence time of the bundle is witnessed by a third party rather than asserted by the producer.
Non-determinism, especially LLM calls, undermines “reproducibility”.Replay re-executes the workflow against the recorded capability results: LLM calls, HTTP fetches, and other effects return their captured responses instead of hitting live services, and --verify checks that the replay reproduces the same output hashes. This proves the recorded run is internally consistent — the recorded inputs deterministically produce the recorded outputs.Replay proves reproducibility given the recorded capability results — it does not prove that the model (or any external service) would return the same thing if called again live. A non-deterministic model is captured, not tamed: the bundle pins what the model said this time, not what the model will say next time. Do not read a passing replay as “the model is deterministic.”
The environment fingerprint is forged.env_fingerprint.json records OS, architecture, and Boruna version, and it is checksummed and covered by bundle_hash like every other file — so it cannot be changed after sealing without detection.The fingerprint is self-reported by the recording process, not hardware-attested. A malicious or misconfigured producer can write any values it likes at seal time; the integrity check only proves those values were not altered afterward, not that they were true. Mitigation direction (not yet implemented): TEE remote attestation, binding the fingerprint to a hardware root of trust that attests the actual code image and platform that ran.

4. Why “detects tampering” needs a footnote

Plain boruna evidence verify gives you internal consistency: it recomputes every checksum and the chain and confirms they agree with the manifest. That catches accidental corruption and unsophisticated edits.

It does not, by itself, catch an attacker who controls the whole bundle, because that attacker can make the manifest agree with their forgery. Two independent, composable checks close this gap; neither is on by default, and each roots trust in something the attacker does not control:

  1. External anchor (--expected-bundle-hash / expected_bundle_hash). You record the bundle_hash out-of-band at seal time — in a separate system the attacker cannot rewrite — and supply it at verify time. Verification then requires the recomputed hash to equal your anchor, not the manifest’s self-reported one. A forged-but-self-consistent bundle fails because its recomputed hash no longer matches the anchor you kept. This is what makes a plaintext bundle genuinely tamper-evident against a motivated attacker.

  2. ed25519 signature under a pinned key (--verify-key / trusted_pubkey, optionally require_signature). The producer signs bundle_hash with an ed25519 key; the verifier pins the expected public key. Trust is rooted in the pinned key: a bundle re-signed with any other key is rejected as signature_untrusted_key. Without pinning, a signature proves only that some key signed — an attacker can substitute their own.

The signature’s meaning is precise: it attests which key sealed these bytes. That is an origin/authenticity claim about the producer, not a truth claim about the content (contrast the malicious-producer row in §3). Non-repudiation follows only to the extent that the key is bound to an accountable identity and is not shared — conditions the bundle format cannot enforce on its own.


5. Mitigation directions (not yet implemented)

The residual gaps in §3 are real. The honest position is that they are known and have known remedies on the roadmap, none of which ship today:

  • Transparency-log anchoring (Rekor-style): witness the bundle_hash in an external append-only log at seal time, giving a third-party- attested earliest-existence timestamp and defeating silent backdating.
  • Keyless / identity-bound signing: bind a signature to a verifiable workload identity for a short-lived credential, reducing the blast radius of a stolen long-lived key.
  • TEE remote attestation: replace the self-reported environment fingerprint with a hardware-attested measurement of the code image and platform that actually executed.

Until these land, treat the corresponding claims conservatively: a bundle proves post-seal integrity and internal replay-consistency, anchored or signed bundles additionally prove origin against a chosen root of trust, and nothing in the bundle proves the producer was honest or that the timestamps are true.


6. See also

  • Evidence Bundles — on-disk layout, hash chain, and the verify / inspect / replay workflow.
  • Runtime Execution Provenance — the provenance category Boruna occupies, and how it relates to SLSA, in-toto, Sigstore, C2PA, and TEE attestation.
  • docs/spec/evidence-bundle-1.0.md — the normative format and integrity contract.

Runtime Execution Provenance

Boruna occupies a provenance category that the established supply-chain and content-authenticity standards leave largely unoccupied: runtime execution provenance — an attested, tamper-evident record of what a specific execution actually did.

The claim this category makes is narrow and concrete:

This specific run executed these steps, made these policy-gated capability calls, under this policy, transforming these recorded inputs into these recorded outputs.

That is a statement about an execution trace, not about a build, an artifact, a signing event, a media file, or a booted code image. The distinction matters because the standards people reach for by reflex all answer a different question.


1. The provenance landscape

Each of these standards is good at what it targets. None of them targets the execution trace of a particular run.

Standard / mechanismWhat it attestsWhat it leaves unoccupied
SLSABuild provenance: that an artifact was produced by a particular build system from particular sources, following a particular process.Says nothing about what happens when the built thing later runs. A SLSA-attested binary can still do anything at runtime.
in-totoSupply-chain step metadata: that each link in a defined software supply chain was performed by an authorized party on declared materials/products.Models the pipeline that assembles software, not the runtime behavior of a workflow execution. The “steps” are build/release steps, not gated capability calls made during one run.
Sigstore / RekorSigning events over artifacts: that a given artifact digest was signed by a given identity, witnessed in an append-only transparency log.Attests the existence of a signature over a blob at a time — not what an execution did. Rekor witnesses that something was signed, not that a run made these calls under this policy.
C2PAContent provenance: the capture/edit history and origin of a media asset (image, audio, video).Concerned with the lineage of content, not with the execution of a program or workflow.
TEE remote attestationCode identity: which enclave/image was measured and booted on attested hardware.Attests what code was loaded and the platform it ran on — not the trace of what that code then did (which steps, which capability calls, which inputs→outputs).

Read down the right-hand column: build provenance, supply-chain step metadata, signing events, content lineage, and code identity are all covered — and the runtime execution trace falls through the gap between them. Boruna’s evidence bundle is aimed squarely at that gap.

These standards are complementary, not competitors. A mature deployment might use SLSA for the binary, TEE attestation for the platform, and Boruna for the execution trace — each answering the question it is built to answer.


2. What Boruna attests

The evidence bundle records the run itself. The load-bearing components (see Evidence Bundles and orchestrator/src/audit/evidence.rs) map directly onto the category claim:

  • These steps executed — the hash-chained audit_log, whose entries record step start/completion and capability calls in order, with each entry chained to the previous (entry_hash = SHA-256(prev_hash || event_json)).
  • Under this policy — policy.json and its policy_hash; the policy snapshot that gated the run is sealed alongside the trace, so a verifier sees the exact rules in force.
  • Made these gated capability calls — capability calls appear in the audit log / event stream; the optional model_invoking_steps.json additionally records which steps transitively reached an LLM capability, so an auditor can see which steps touched a model without re-analyzing sources.
  • Transforming these inputs into these outputs — per-step outputs under outputs/<step>/<name>.json, each SHA-256-checksummed in file_checksums and thereby committed to by bundle_hash.
  • With declared purpose — the optional intents.json records the per-step declared intent (what each step was authorized to do), captured as replay-verified evidence alongside what it actually did.

All of these are covered by the same integrity contract: on-disk checksums, an unbroken audit-log chain, and a bundle_hash over the manifest. The record is therefore tamper-evident in exactly the sense defined in the Evidence Bundle Threat Model — and, as that document is careful to state, tamper-evidence is not tamper-proofing, and a sealed trace attests what was recorded, not that the producer recorded honestly.


3. Interoperability with the standards

Occupying a distinct category does not mean living apart from the ecosystem. The intent is for a Boruna execution record to slot into existing supply-chain and attestation tooling rather than replace it.

To that end, an in-toto / DSSE emission is available via boruna evidence attest <bundle-dir>: it exports the bundle’s core attestation (run identity, workflow_hash, policy_hash, audit_log_hash, output digests) as an in-toto Statement (predicateType: https://boruna.dev/runtime-provenance/v1) wrapped in a DSSE envelope signed with the bundle’s ed25519 key. That makes the execution-provenance claim consumable by the same tooling that already ingests SLSA and in-toto attestations — for example, a transparency log or a policy engine that gates on attestations. The predicate schema is specified in docs/spec/runtime-provenance-predicate-1.0.md.

Status note. The in-toto/DSSE emitter is implemented (boruna evidence attest, --verify to check the envelope). It is additive — the native, authoritative format remains the evidence bundle described in docs/spec/evidence-bundle-1.0.md. Live interoperability with a specific cosign/in-toto-verify binary is spec-conformant but not yet end-to-end verified (the DSSE keyid is a raw-hex ed25519 key that a consumer must bridge to PEM SPKI); see the predicate spec’s compatibility notes.

The relationship is layered, not overlapping:

  • SLSA / in-toto attest how the software (and, via DSSE, other attestations) was produced and assembled.
  • Sigstore / Rekor can witness signatures — including, once emitted, a DSSE-wrapped Boruna attestation — in an append-only log.
  • TEE attestation can vouch for the platform and code image that ran the Boruna engine.
  • Boruna attests the execution trace that happened on top of all of the above.

Each layer roots a different claim; together they compose into a story that no single standard tells alone.


4. See also

  • Evidence Bundles — the concrete artifact that carries the execution-provenance record.
  • Evidence Bundle Threat Model — precisely what the record proves and does not prove (tamper-evidence vs. tamper-proofing vs. non-repudiation).
  • docs/spec/evidence-bundle-1.0.md — the normative on-disk format and integrity contract.

Bundle Storage

Evidence bundles are written to local disk by the orchestrator (see Evidence Bundles). The Bundle Storage layer is an optional second-stage that copies the finalized bundle to a durable remote backend after the local write succeeds.

The local bundle is always the authoritative record. Remote storage is additive durability — operators choose to enable it for compliance, archival, or cross-region disaster recovery.

Backends

SchemeBackendFeature flagGuide
local:<root>Local filesystemalways available(no separate guide — see CLI help)
s3://<bucket>[/<prefix>]AWS S3 / S3-compatible (MinIO, R2, B2)s3bundle-storage-s3.md
gs://<bucket>[/<prefix>]Google Cloud Storagegcsbundle-storage-gcs.md
azblob://<account>/<container>[/<prefix>]Azure Blob Storageazurebundle-storage-azure.md

All four schemes share the same trait, the same dispatcher, and the same failure contract — everything below applies to every backend.

How it works

Pass the URI to boruna workflow run --record:

boruna workflow run my-workflow \
  --policy allow-all \
  --record \
  --bundle-storage s3://my-audit-bucket/prod

After the run finalizes locally, the orchestrator:

  1. Resolves the URI to a BundleStorage adapter via from_uri.
  2. Calls storage.put(run_id, bundle_dir).
  3. Records the returned StorageRef in the run output as storage_ref.

The BORUNA_BUNDLE_STORAGE env var is the standard way to set this once for an entire deployment.

Failure contract

A storage failure never masks a successful workflow.

If the remote backend is unreachable, returns auth errors, or hits any other backend failure:

  • The orchestrator logs a warning to stderr.
  • The local bundle remains on disk and remains the authoritative record.
  • The workflow exits successfully (the remote copy is not required for the run to be considered complete).
  • storage_ref is omitted from the output.

Operators who require remote durability should monitor the warning output and re-run evidence verify against the local bundle to confirm it is still on disk before pruning it.

Reading bundles back

Once you have a storage_ref, materialize the bundle into a local cache directory:

#![allow(unused)]
fn main() {
use boruna_orchestrator::audit::storage::{from_uri, StorageRef};

let storage = from_uri(Some("s3://my-audit-bucket/prod"))?.unwrap();
let local_dir = storage.get(&StorageRef("s3://my-audit-bucket/prod/<run-id>".into()))?;
boruna_orchestrator::audit::verify_bundle(&local_dir)?;
}

The cache directory defaults to <temp>/boruna-bundle-cache and is overridable via BORUNA_BUNDLE_CACHE. It is shared across all remote adapters; if you operate multi-cloud and care about cross-bucket consistency, set per-bucket cache directories yourself.

OFF-feature behavior

When you build the binary without the feature for a remote scheme, the corresponding URI rejects at parse time with an actionable message that points at the feature flag:

$ boruna workflow run --bundle-storage s3://b/p ...
warning: --bundle-storage URI invalid: s3://b/p requires the `s3` feature; rebuild with `--features boruna-cli/s3`

This is a deliberate design choice: silently ignoring the URI would create an audit gap (the operator believes their bundle is going to S3 but it is not). The warning is loud, the build directive is explicit, and the workflow still completes successfully against local storage.

Determinism contract

The storage_ref returned by put is operational metadata — it does not feed any audit-log hash or replay comparison. The bundle’s bundle_hash and audit_log_hash are computed from the local manifest before remote upload begins; where the bundle is also stored does not affect those hashes.

This means: a bundle uploaded to S3 yesterday and to GCS today carries the same bundle_hash. The two storage_refs differ; the underlying audit content is identical.

Stable error taxonomy

BundleStorage::put / get / list return StorageError. The Backend { kind, msg } variant carries a stable per-adapter kind string so integrators can branch on retry semantics:

Adapterkind taxonomy
S3s3.transient, s3.permanent, s3.runtime, s3.unexpected_key
GCSgcs.transient, gcs.permanent, gcs.runtime, gcs.unexpected_key
Azureazure.transient, azure.permanent, azure.runtime, azure.unexpected_key

StorageError is #[non_exhaustive], so adding a new variant in a future release is additive (existing match arms keep compiling). Likewise, new kind strings are additive — integrators switching on kind should treat unknown values as transient (retryable) by default.

See also

  • Evidence Bundles — what’s inside a bundle and how the audit chain works
  • The per-adapter operator guides linked in the table above
  • API reference: boruna_orchestrator::audit::storage — the trait, dispatcher, and LocalFs adapter
  • API reference: boruna_orchestrator::audit::storage_{s3,gcs,azure} — the per-provider implementations (each behind its own feature flag)

Compliance Evidence

Evidence Bundles

Every workflow run with --record produces a self-contained evidence bundle — a directory of artifacts sufficient for compliance audit.

Bundle Contents

FilePurpose
manifest.jsonBundle metadata, checksums, environment info
workflow.jsonExact workflow definition used
policy.jsonExact policy applied
audit_log.jsonHash-chained log of all decisions and events
env_fingerprint.jsonRuntime environment (Boruna version, OS, arch)
outputs/<step>/<name>.jsonPer-step output data

Manifest Fields

{
  "schema_version": 1,
  "run_id": "run-my-workflow-20260221T120000",
  "workflow_name": "my-workflow",
  "workflow_hash": "<sha256>",
  "policy_hash": "<sha256>",
  "audit_log_hash": "<sha256>",
  "file_checksums": {
    "workflow.json": "<sha256>",
    "policy.json": "<sha256>",
    "audit_log.json": "<sha256>"
  },
  "env_fingerprint": {
    "boruna_version": "0.1.0",
    "rust_version": "...",
    "os": "linux",
    "arch": "x86_64",
    "hostname": "worker-01"
  },
  "started_at": "2026-02-21T12:00:00Z",
  "completed_at": "2026-02-21T12:00:01Z",
  "bundle_hash": "<sha256>"
}

Verification

boruna evidence verify <bundle-dir>

Verification checks:

  1. All files listed in file_checksums exist and match their SHA-256 hashes
  2. Audit log hash-chain is intact (each entry’s hash includes the previous)
  3. Audit log hash matches the manifest’s audit_log_hash
  4. Required files are present (manifest, workflow, policy, audit log, fingerprint)

Audit Log

The audit log is a JSON array of hash-chained entries:

[
  {
    "sequence": 0,
    "prev_hash": "0000000000000000000000000000000000000000000000000000000000000000",
    "event": { "WorkflowStarted": { "workflow_hash": "...", "policy_hash": "..." } },
    "entry_hash": "<sha256>"
  },
  {
    "sequence": 1,
    "prev_hash": "<hash of entry 0>",
    "event": { "StepCompleted": { "step_id": "fetch", "output_hash": "...", "duration_ms": 42 } },
    "entry_hash": "<sha256>"
  }
]

Tamper Detection

Each entry’s hash = SHA-256(sequence + prev_hash + event_json). Modifying any entry invalidates all subsequent hashes.

Determinism Proof

To prove determinism:

  1. Run the same workflow twice with --record
  2. Compare evidence bundles: file checksums should be identical (excluding timestamps)
  3. Audit log event hashes should match for the same events

What This Proves

For compliance auditors, an evidence bundle demonstrates:

  • What ran: Exact workflow definition and policy
  • When it ran: Timestamps and run ID
  • Where it ran: Environment fingerprint
  • What happened: Complete audit trail of every step, decision, and output
  • Integrity: SHA-256 hash chain prevents undetected modification

LLM Integration Guide

Boruna does not ship a default LLM handler. The supported integration model is Bring Your Own Handler (BYOH) — you implement the CapabilityHandler trait against your provider of choice, and pass the handler into CapabilityGateway::with_handler at workflow run time.

This document explains why, what the contract looks like, and how to wire up handlers for common providers (OpenAI, Anthropic, vLLM, Ollama).

Why BYOH?

Three reasons.

  1. Provider churn doesn’t destabilize Boruna. OpenAI’s API changes; so does Anthropic’s; so does the long tail. If a default handler shipped in core, every provider release would risk a Boruna patch release. With BYOH, your handler tracks your provider on your release cadence.

  2. API key management belongs in your application. Boruna is a deterministic execution chassis. Secret loading, rotation, observability, billing attribution — all of these are concerns that vary wildly per organization. Pushing them into core would constrain integrators (e.g. FleetQ) who already have their own conventions.

  3. Most production integrators already have an LLM client. The platform’s primary integrators run their own LLM infrastructure (vLLM clusters, OpenRouter proxies, custom routing). Shipping a default handler would just add another thing for them to override.

A “convenience CLI handler” was considered and rejected — it would lock the project into provider compatibility commitments without serving the production integrators it’s primarily meant for.

The contract

CapabilityHandler is a single-method trait in boruna-vm::capability_gateway:

#![allow(unused)]
fn main() {
pub trait CapabilityHandler: Send {
    fn handle(&mut self, cap: &Capability, args: &[Value]) -> Result<Value, String>;
}
}

For LLM calls, the relevant capability is Capability::LlmCall. The first argument is conventionally the prompt (a Value::String); subsequent arguments are provider-specific. The return is the LLM’s response, also conventionally Value::String or Value::Map { "content": ..., "finish_reason": ... }.

The handler runs inside the VM’s step execution. It has Send bound but no Sync — each VM instance owns its own handler. In the concurrent execution path (--concurrency > 1), each worker constructs its own handler.

Minimal example

#![allow(unused)]
fn main() {
use boruna_bytecode::{Capability, Value};
use boruna_vm::capability_gateway::CapabilityHandler;

pub struct OpenAiHandler {
    api_key: String,
    client: ureq::Agent,
}

impl OpenAiHandler {
    pub fn from_env() -> Result<Self, String> {
        let api_key = std::env::var("OPENAI_API_KEY")
            .map_err(|_| "OPENAI_API_KEY not set".to_string())?;
        Ok(Self {
            api_key,
            client: ureq::AgentBuilder::new()
                .timeout(std::time::Duration::from_secs(60))
                .build(),
        })
    }
}

impl CapabilityHandler for OpenAiHandler {
    fn handle(&mut self, cap: &Capability, args: &[Value]) -> Result<Value, String> {
        match cap {
            Capability::LlmCall => self.handle_llm_call(args),
            // Delegate everything else to the default mock handler.
            _ => boruna_vm::capability_gateway::MockHandler.handle(cap, args),
        }
    }
}

impl OpenAiHandler {
    fn handle_llm_call(&mut self, args: &[Value]) -> Result<Value, String> {
        let prompt = args
            .first()
            .and_then(|v| match v {
                Value::String(s) => Some(s.as_str()),
                _ => None,
            })
            .ok_or_else(|| "llm.call: first arg must be a String prompt".to_string())?;

        let body = serde_json::json!({
            "model": "gpt-4o-mini",
            "messages": [{"role": "user", "content": prompt}],
        });

        let resp = self
            .client
            .post("https://api.openai.com/v1/chat/completions")
            .set("Authorization", &format!("Bearer {}", self.api_key))
            .set("Content-Type", "application/json")
            .send_json(body)
            .map_err(|e| format!("openai request failed: {e}"))?;

        let json: serde_json::Value = resp
            .into_json()
            .map_err(|e| format!("openai response parse: {e}"))?;
        let content = json
            .get("choices")
            .and_then(|c| c.get(0))
            .and_then(|c| c.get("message"))
            .and_then(|m| m.get("content"))
            .and_then(|c| c.as_str())
            .ok_or_else(|| "openai response missing choices[0].message.content".to_string())?;

        Ok(Value::String(content.to_string()))
    }
}
}

A complete reference handler lives in examples/llm_handlers/openai/. The directory contains handler.rs (a self-contained module you can copy into your integrator crate) and a README documenting the per-provider gotchas.

Provider variants

Reference handlers ship for the most common providers (post1-T-1.2). Each is a copy-and-tweak template, not a compiled crate Boruna pulls in:

ProviderUse whenPath
OpenAIapi.openai.comexamples/llm_handlers/openai/
Anthropicapi.anthropic.comexamples/llm_handlers/anthropic/
Ollamalocal development, air-gapped CIexamples/llm_handlers/ollama/
vLLM (and OpenAI-compatible proxies)self-hosted vLLM, OpenRouter, Together, Groq, LiteLLMexamples/llm_handlers/vllm/
AWS BedrockBedrock-hosted Claude / Llama / Titan / Mistral / Cohereexamples/llm_handlers/bedrock/

The umbrella examples/llm_handlers/README.md gives the cross-provider overview and links each handler’s gotchas (auth header, response path, determinism options).

A documented providers.toml.example schema and a router_setup.rs reference parser show one convention for declaring your provider lineup in config rather than code. The format is a starting point — Boruna does not parse it.

For multi-provider routing, use the built-in LlmRouterHandler (covered next), or have your handler switch on a model argument (passed as args[1]) or on a thread-local context.

LlmRouterHandler — built-in multi-provider dispatch (sprint 0.4-S13)

If you have multiple providers wired in (OpenAI + Anthropic + a local Ollama, say) and don’t want to write your own dispatch logic, Boruna ships a LlmRouterHandler in boruna-vm::capability_gateway. It takes a registry of provider handlers keyed by name and dispatches each Capability::LlmCall based on a provider/model prefix in args[1]:

#![allow(unused)]
fn main() {
use std::collections::BTreeMap;
use boruna_vm::capability_gateway::{
    CapabilityHandler, LlmRouterHandler, MockHandler,
};

let mut providers: BTreeMap<String, Box<dyn CapabilityHandler>> = BTreeMap::new();
providers.insert("openai".into(), Box::new(my_openai_handler));
providers.insert("anthropic".into(), Box::new(my_anthropic_handler));
providers.insert("ollama".into(), Box::new(my_ollama_handler));

// Non-LLM calls pass through to the fallback (typically MockHandler
// in tests; in production you'd plug a HttpHandler etc).
let router = LlmRouterHandler::new(providers, Box::new(MockHandler));

// Pass `router` into `CapabilityGateway::with_handler` like any
// other CapabilityHandler.
}

.ax callers then write:

let response = llm_call("Summarize:", "openai/gpt-4")
let response2 = llm_call("Translate:", "anthropic/claude-3-5-sonnet-20241022")

The full model string (including the provider/ prefix) is forwarded to the provider’s handler unchanged, so providers can use the prefix internally (e.g. for billing tags).

The router does not impose provider compatibility commitments on core — it’s pure routing logic. You still bring your own per-provider handler implementation. boruna-vm ships zero provider HTTP code.

Determinism considerations

LLM calls are inherently non-deterministic at the model layer (sampling temperature, randomness, model version drift). Three things keep Boruna’s determinism contract intact even with non-deterministic handlers:

  1. output_hash reflects what was actually returned. If the LLM returned different text on two runs, the output_hash differs and downstream replay-comparison surfaces the divergence.
  2. net.fetch record-replay (sprint 0.5-S7) lets you record an LLM session and replay deterministically — useful for testing and audit.
  3. Persisted output_json (sprint 0.3-S2b) means a successful LLM call’s output survives process restarts. Resume picks up the recorded value rather than re-calling the LLM.

For workflows that need bit-identical replay, set temperature: 0 and a pinned model version, then use --record-net-to to capture the network transactions for future replay.

Capability policy

Steps that call LLMs must declare the capability in the workflow definition:

{
  "steps": {
    "summarize": {
      "kind": "source",
      "source": "steps/summarize.ax",
      "capabilities": ["llm.call"],
      "budget": { "max_calls": 3 }
    }
  }
}

The budget.max_calls field caps how many llm.call invocations the step is allowed. Exceeding the budget returns a CapabilityBudgetExceeded runtime error. Per-step budgeting is enforced by CapabilityGateway regardless of which handler is plugged in.

Testing your handler

Two patterns:

  1. Mock at the trait level — pass a stub CapabilityHandler that returns canned responses. Fast, doesn’t hit the network. Use this for the bulk of your test suite.
  2. Record + replay — run once with the real handler and --record-net-to <tape>, then in tests run with --replay-net-from <tape>. Captures the actual provider response shape but stays offline. Recommended for integration tests and CI.

Where to look in the code

  • crates/llmvm/src/capability_gateway.rs — CapabilityHandler trait + MockHandler (default).
  • crates/llm-effect/ — higher-level prompt, cache, context primitives (provider-agnostic).
  • examples/llm_handlers/ — reference handler implementations.

What this guide is NOT

  • Not a tutorial on prompt engineering. See crates/llm-effect/’s prompt module for the prompt-building primitives Boruna ships.
  • Not a recommendation of a specific provider. Pick what your organization standardizes on.
  • Not a streaming-API guide. Boruna’s capability calls are synchronous; streaming responses must be collected to a single value before returning.

Model Evaluation Framework

boruna workflow eval runs the same workflow against two LLM provider configurations and compares the resulting evidence bundles.

Usage

boruna workflow eval examples/workflows/llm_code_review \
  --providers-a anthropic.json \
  --providers-b ollama.json \
  --runs 3

Provider config format

{
  "llm.call": {
    "provider": "anthropic",
    "model": "claude-3-5-sonnet-20241022",
    "api_key_env": "ANTHROPIC_API_KEY"
  }
}

Supported providers: anthropic, ollama, openai_compat, deny, passthrough.

See docs/guides/llm-integration.md for full BYOH handler wiring.

Output

The command prints a comparison table:

Provider A (anthropic): 3/3 runs succeeded (100%), mean 1420ms
Provider B (ollama):    3/3 runs succeeded (100%), mean 890ms

Step                     A identical   B identical   A vs B agree
----------------------------------------------------------------------
analyze                  yes           yes           no  (different)
report                   yes           yes           yes (identical)

Use --json for machine-readable output suitable for CI pipelines.

Options

FlagDescription
--providers-a <file>First provider config JSON
--providers-b <file>Second provider config JSON
--runs NNumber of runs per provider (default: 1)
--data-dir <dir>Directory for evidence bundles (default: .boruna/data)
--jsonMachine-readable JSON output

Evidence bundles

Each run writes a full evidence bundle under <data-dir>/model-eval/<provider-name>/run_N/evidence/. These can be inspected with boruna evidence inspect or diffed with boruna evidence diff.

Notes on live LLM comparison

The eval command validates the config and logs the provider that would be used, then runs the workflow in demo mode. Live provider comparison (actual LLM API calls) requires the boruna-vm/http feature and wired BYOH handlers — see docs/guides/llm-integration.md.

Testing Guide

TestHarness

The primary testing tool. No host UI required.

#![allow(unused)]
fn main() {
use boruna_framework::testing::TestHarness;
use boruna_framework::runtime::AppMessage;
use boruna_bytecode::Value;

let mut harness = TestHarness::from_source(SOURCE)?;
}

Send Messages

#![allow(unused)]
fn main() {
let (state, effects) = harness.send(
    AppMessage::new("increment", Value::Int(0))
)?;
}

Simulate Sequences

#![allow(unused)]
fn main() {
let final_state = harness.simulate(vec![
    AppMessage::new("add", Value::Int(0)),
    AppMessage::new("add", Value::Int(0)),
    AppMessage::new("complete", Value::Int(0)),
])?;
}

Assertions

#![allow(unused)]
fn main() {
// Check a specific field by index
harness.assert_state_field(0, &Value::Int(3))?;

// Check full state equality
harness.assert_state(&expected_value)?;

// Check effects from last cycle
harness.assert_effects(&["http_request"])?;
}

Snapshots

#![allow(unused)]
fn main() {
let json = harness.snapshot();  // JSON string of current state
}

Time Travel

#![allow(unused)]
fn main() {
harness.rewind(0)?;  // Go back to init state
}

Replay Verification

#![allow(unused)]
fn main() {
let messages = vec![
    AppMessage::new("increment", Value::Int(0)),
    AppMessage::new("increment", Value::Int(0)),
];
for msg in &messages {
    harness.send(msg.clone())?;
}

// Replay same messages on fresh runtime, verify identical states
let identical = harness.replay_verify(SOURCE, messages)?;
assert!(identical);
}

Golden Tests

Golden tests verify determinism. Run the same messages twice, compare:

#![allow(unused)]
fn main() {
let mut h1 = TestHarness::from_source(SOURCE)?;
let mut h2 = TestHarness::from_source(SOURCE)?;

for msg in &messages {
    h1.send(msg.clone())?;
    h2.send(msg.clone())?;
}

assert_eq!(h1.state(), h2.state());
assert_eq!(h1.snapshot(), h2.snapshot());
}

CLI Testing

# Validate app protocol
boruna framework validate my_app.ax

# Send messages and see state
boruna framework test my_app.ax -m "add:0,add:0,complete:0"

# Inspect state as JSON
boruna framework inspect-state my_app.ax -m "add:0,add:0"

# Step-by-step simulation
boruna framework simulate my_app.ax "add:0,add:0,complete:0"

# Machine-readable diagnostics
boruna framework diag my_app.ax -m "add:0,add:0"

# Stable trace hash (for CI comparison)
boruna framework trace-hash my_app.ax -m "add:0,add:0"

# App contract summary
boruna framework inspect my_app.ax --json

Message Format

CLI messages use tag:payload format:

  • increment:0 → tag=“increment”, payload=Int(0)
  • fetch:https://example.com → tag=“fetch”, payload=String(“https://example.com”)
  • reset → tag=“reset”, payload=Int(0) (default)

Comma-separated for sequences: add:0,add:0,complete:0

Writing Test Apps

Include fn main() -> Int for standalone execution:

fn main() -> Int {
    let s0: State = init()
    let r1: UpdateResult = update(s0, Msg { tag: "add", payload: 0 })
    let r2: UpdateResult = update(r1.state, Msg { tag: "add", payload: 0 })
    r2.state.total
}

Run directly: boruna run my_app.ax

Trace → Regression Tests + Minimizer

Overview

The trace2tests module converts runtime execution traces into deterministic regression tests. It also provides a delta-debugging minimizer that shrinks failing traces to minimal reproducing sequences.

Trace Schema

Version 1, stable JSON format.

{
  "version": 1,
  "source_file": "path/to/app.ax",
  "source_hash": "sha256:<hex>",
  "cycles": [
    {
      "cycle": 1,
      "message": { "tag": "increment", "payload": {"Int": 0} },
      "state_before_hash": "sha256:<hex>",
      "state_after_hash": "sha256:<hex>",
      "state_after": {"Record": {"type_id": 0, "fields": [{"Int": 1}]}},
      "effects": [
        { "kind": "http_request", "payload_hash": "sha256:<hex>", "callback_tag": "on_response" }
      ],
      "ui_tree_hash": "sha256:<hex>"
    }
  ],
  "final_state_hash": "sha256:<hex>",
  "trace_hash": "sha256:<hex>"
}

Fields

FieldTypeDescription
versionu32Schema version (always 1)
source_filestringPath to source .ax file
source_hashstringSHA-256 of source text
cyclesarrayOrdered cycle records
final_state_hashstringSHA-256 of final state
trace_hashstringSHA-256 of canonical fingerprint

Hashing

All hashes use SHA-256 of canonical JSON serialization:

  • Values are serialized via serde (deterministic for BTreeMap)
  • The trace fingerprint concatenates all cycle data in stable format
  • Same inputs always produce identical hashes

Test Spec Format

Generated test specifications are self-contained JSON:

{
  "version": 1,
  "name": "counter_regression",
  "source_file": "examples/counter.ax",
  "source_hash": "sha256:<hex>",
  "messages": [
    { "tag": "increment", "payload": {"Int": 0} },
    { "tag": "decrement", "payload": {"Int": 0} }
  ],
  "assertions": [
    { "kind": "final_state_hash", "expected": "sha256:<hex>", "description": "..." },
    { "kind": "trace_hash", "expected": "sha256:<hex>", "description": "..." },
    { "kind": "cycle_count", "expected": "2", "description": "..." }
  ]
}

Assertion Kinds

KindDescription
final_state_hashSHA-256 of final state matches
trace_hashSHA-256 of full trace fingerprint matches
cycle_countNumber of cycles matches

Delta Debugging Minimizer

Implements the ddmin algorithm to shrink failing message sequences:

  1. Chunk removal: Split into n chunks, try removing each
  2. Granularity increase: If no chunk removal works, try finer splits
  3. 1-minimal pass: Try removing each individual message
  4. Result is guaranteed 1-minimal (removing any single message stops the failure)

Predicates

Built-in predicates:

  • panic: Failure = runtime error during message processing
  • State mismatch: Failure = final state hash differs from expected

External predicates: Any command that receives a temp trace file path and returns non-zero on failure.

CLI Usage

Record

boruna trace2tests record <file.ax> --messages "tag:payload,..." --out trace.json

Generate

boruna trace2tests generate --trace trace.json --out test_spec.json [--name test_name]

Run

boruna trace2tests run --spec test_spec.json [--source app.ax]

Minimize

boruna trace2tests minimize --trace trace.json --source app.ax [--predicate panic]
boruna trace2tests minimize --trace trace.json --source app.ax --predicate "my_check.sh"

Determinism Guarantees

  • Same source + same messages = identical trace hash
  • Generated tests are deterministic regression gates
  • Minimizer produces deterministic output (same input → same minimal trace)
  • All hashes use SHA-256 with canonical serialization

Integration

  • Trace files are compatible with the framework’s CycleRecord format
  • Test specs can be version-controlled alongside source
  • Minimized traces export as regression tests via generate
  • The full pipeline: record → minimize → generate → run

App Template

Canonical File Layout

my_app/
  my_app.ax       # Single source file

Minimal App Skeleton

// my_app — Boruna Framework App

type State { value: Int }
type Msg { tag: String, payload: Int }
type Effect { kind: String, payload: String, callback_tag: String }
type UpdateResult { state: State, effects: List<Effect> }
type UINode { tag: String, text: String }

fn init() -> State {
    State { value: 0 }
}

fn update(state: State, msg: Msg) -> UpdateResult {
    UpdateResult {
        state: State { value: state.value + msg.payload },
        effects: [],
    }
}

fn view(state: State) -> UINode {
    UINode { tag: "text", text: "value" }
}

fn main() -> Int {
    let s: State = init()
    s.value
}

Required Types

TypeFields
StateYour app state. Must be a record.
Msgtag: String, payload: <T>. Tag dispatches logic.
Effectkind: String, payload: String, callback_tag: String
UpdateResultstate: State, effects: List<Effect>
UINodetag: String + any additional fields

Required Functions

FunctionSignaturePure?
init()() -> StateNo
update()(State, Msg) -> UpdateResultYes
view()(State) -> UINodeYes

Optional Functions

FunctionSignaturePurpose
policies()() -> PolicySetDeclare allowed capabilities
main()() -> IntStandalone test entry point

Create From CLI

boruna framework new my_app
boruna framework validate my_app/my_app.ax
boruna run my_app/my_app.ax

Policy Template

type PolicySet { capabilities: List<String>, max_effects: Int, max_steps: Int }

fn policies() -> PolicySet {
    PolicySet {
        capabilities: ["net.fetch", "time.now"],
        max_effects: 10,
        max_steps: 1000000,
    }
}

Effects Guide

Overview

Effects are declarative descriptions of side effects. update() never performs IO directly. Instead, it returns a list of effects. The framework runtime executes them via the capability gateway.

Effect Structure

type Effect { kind: String, payload: String, callback_tag: String }
  • kind — which effect to execute (see table below)
  • payload — data for the effect (URL, query, path, etc.)
  • callback_tag — message tag for the result delivery

Built-in Effect Kinds

KindCapabilityDescription
http_requestnet.fetchHTTP GET/POST request
db_querydb.queryDatabase query
fs_readfs.readRead file
fs_writefs.writeWrite file
timertime.nowGet current time
randomrandomGet random value
spawn_actorspawnSpawn child actor
emit_uiui.renderEmit UI tree to host

Returning Effects From update()

fn update(state: State, msg: Msg) -> UpdateResult {
    if msg.tag == "fetch" {
        UpdateResult {
            state: state,
            effects: [
                Effect {
                    kind: "http_request",
                    payload: "https://api.example.com/data",
                    callback_tag: "data_received",
                },
            ],
        }
    } else {
        UpdateResult { state: state, effects: [] }
    }
}

Effect Lifecycle

  1. update() returns UpdateResult { state, effects }.
  2. Framework validates effects against the policy.
  3. Framework (or host) executes each effect via capability gateway.
  4. Effect results are delivered as new messages with callback_tag as the tag.
  5. update() handles the callback message in the next cycle.

Multiple Effects Per Cycle

Return multiple effects in the list. They execute in order.

effects: [
    Effect { kind: "http_request", payload: "url1", callback_tag: "result" },
    Effect { kind: "http_request", payload: "url2", callback_tag: "result" },
    Effect { kind: "http_request", payload: "url3", callback_tag: "result" },
]

Policy Constraints

Effects are checked against the app’s PolicySet:

  • Only listed capabilities are allowed.
  • max_effects_per_cycle limits how many effects per update.
  • Violations produce FrameworkError::PolicyViolation.

Determinism

Effects themselves are deterministic data. Their execution results are logged by the capability gateway. Replay substitutes recorded results, making the entire execution deterministic.

Actors Guide

Status

Actor integration in the framework is partial. The VM has an ActorSystem with basic spawn/send/receive mechanics. The framework does not yet wire actors into the App protocol runtime.

Architecture

Parent App
  ├── init() → State
  ├── update(state, msg) → UpdateResult
  │     └── effects: [Effect { kind: "spawn_actor", ... }]
  ├── view(state) → UINode
  └── Child Actor (same App protocol)
        ├── init() → State
        ├── update(state, msg) → UpdateResult
        └── view(state) → UINode

Spawning Actors

Return a spawn_actor effect from update():

Effect { kind: "spawn_actor", payload: "child_module", callback_tag: "child_spawned" }

The framework runtime will:

  1. Compile and initialize the child module.
  2. Assign it an actor ID.
  3. Deliver a child_spawned message to the parent with the actor ID.

Message Routing

Messages between actors use the tag routing convention:

  • Parent → Child: effect with kind: "send_to_actor" (not yet implemented)
  • Child → Parent: effect with kind: "emit_ui" or framework routing

Supervision

If a child actor crashes (runtime error), the parent receives an error message:

  • Tag: "actor_error"
  • Payload: error description

Scheduling

In single-actor mode: FIFO message processing, deterministic. In multi-actor mode: round-robin scheduling, deterministic order in debug mode.

Current Limitations

  1. Actor spawning is not executed by the framework runtime (only parsed as effects).
  2. No inter-actor message routing at the framework level.
  3. No supervision tree implementation.
  4. The VM’s ActorSystem exists but is not integrated with AppRuntime.

VM-Level Actor API

The VM provides these opcodes for actors:

  • SpawnActor(func_idx) — spawn actor from function
  • SendMsg — send message to actor ID
  • ReceiveMsg — block for incoming message

These are available in bytecode but not yet connected to the framework’s effect-based actor model.

Operations Guide

Running Workflows

Validate

boruna workflow validate <workflow-dir>

Checks DAG structure, step references, input/output wiring, and cycle detection.

Run

boruna workflow run <workflow-dir> --policy allow-all
boruna workflow run <workflow-dir> --policy deny-all
boruna workflow run <workflow-dir> --policy path/to/policy.json

Run with Evidence Recording

boruna workflow run <workflow-dir> --policy allow-all --record
boruna workflow run <workflow-dir> --record --evidence-dir ./evidence

Produces an evidence bundle directory containing:

  • manifest.json — bundle metadata with checksums
  • workflow.json — workflow definition snapshot
  • policy.json — policy snapshot
  • audit_log.json — hash-chained audit entries
  • env_fingerprint.json — runtime environment info
  • outputs/<step_id>/result.json — per-step output data

Verifying Evidence

boruna evidence verify <bundle-dir>

Checks:

  • File checksums match manifest
  • Audit log chain integrity
  • Required files present
boruna evidence inspect <bundle-dir>
boruna evidence inspect <bundle-dir> --json

Displays bundle manifest details.

Replay

For single-file execution:

boruna run app.ax --record trace.json
boruna replay app.axbc trace.json

For workflow-level replay, re-run the workflow in mock mode using recorded outputs.

Observability

Current

  • CLI output shows per-step status, duration, and errors
  • Evidence bundles capture all execution details
  • Audit log provides ordered event history

Planned (Gap)

  • Structured JSON logging via tracing crate
  • Metrics (latency per step, cache hit rate, budget consumption)
  • OpenTelemetry trace export

CI Integration

# Validate all example workflows
for dir in examples/workflows/*/; do
  boruna workflow validate "$dir"
done

# Run in mock mode and verify
boruna workflow run examples/workflows/llm_code_review --record
boruna evidence verify examples/workflows/llm_code_review/evidence/<run-id>

Deployment

Boruna is a statically-linked Rust binary. Deploy by copying the binary to the target system.

cargo build --release --bin boruna
# Binary at: target/release/boruna

Daemon/service mode is documented as a P2 gap in ENTERPRISE_GAPS.md.

Language Server (LSP) Setup

boruna-lsp implements the Language Server Protocol for .ax files.

Build

cargo build --bin boruna-lsp

VS Code

Install the official LSP client extension (coming soon), or configure manually in .vscode/settings.json:

{
  "boruna.lsp.serverPath": "/path/to/boruna-lsp"
}

Or use the generic vscode-languageclient setup with any LSP-capable extension.

Neovim (nvim-lspconfig)

require('lspconfig').boruna_lsp.setup {
  cmd = { '/path/to/boruna-lsp' },
  filetypes = { 'ax' },
  root_dir = require('lspconfig.util').root_pattern('Cargo.toml', '.git'),
}

Features

FeatureStatus
Real-time diagnostics (errors, warnings)✓
Keyword and built-in completion✓
Document formatting (boruna fmt)✓
Hover (function signatures)✓
Go-to-definitionplanned
Rename symbolplanned

Migration tooling (beta)

Status: beta (sprint W5-C). API and on-disk shapes may shift in the 0.6 line; the public CLI invocation will remain stable.

boruna migrate upgrades pre-1.0 artifacts to the format the current release expects. It is the supported way to bring legacy evidence bundles and workflow definitions forward as Boruna’s on-disk contracts evolve toward 1.0.

When to use it

SymptomMigratorNotes
boruna evidence inspect rejects a v0.5.0-or-earlier bundle for missing bundle.jsonevidence-bundleSynthesizes a bundle.json summary; cannot reconstruct missing checksums (see “What it does NOT do” below).
boruna workflow validate rejects a workflow.json with “missing schema_version”workflow-jsonAdds "schema_version": 1.

If you are starting fresh with 0.6.0+, you do not need this tool — new artifacts are produced in the current format.

CLI shape

boruna migrate <kind> <path> [--from <version>] [--to <version>] [--dry-run] [--in-place]
FlagPurpose
<kind>evidence-bundle or workflow-json.
<path>Bundle directory (for evidence-bundle) or workflow.json file (for workflow-json).
--fromSource version of the input. Optional; the migrator infers from contents when absent.
--toTarget version. Defaults to current (latest stable). The beta also accepts 1, 0.6.0, 1.0.0.
--dry-runReport what WOULD change without touching disk. Always run this first.
--in-placeRewrite the input artifact directly. Default is to write a <path>.migrated sibling.

Always do this before shipping a migrated artifact into a downstream pipeline:

# 1. Take a backup. The migrator does not snapshot for you.
cp -R bundles/run-2026-01-01-abc bundles/run-2026-01-01-abc.bak

# 2. Dry-run to see the planned change.
boruna migrate evidence-bundle bundles/run-2026-01-01-abc --dry-run

# 3. Apply, writing a sibling file you can diff.
boruna migrate evidence-bundle bundles/run-2026-01-01-abc

# 4. Review bundle.json.migrated; once happy, swap in place.
boruna migrate evidence-bundle bundles/run-2026-01-01-abc --in-place

For workflow.json the same pattern applies:

boruna migrate workflow-json examples/workflows/legacy/workflow.json --dry-run
boruna migrate workflow-json examples/workflows/legacy/workflow.json --in-place

What each migrator does

evidence-bundle

For a bundle directory:

  • If bundle.json already exists, the migrator reports a no-op.
  • Otherwise, it synthesizes a bundle.json summary containing format_version, boruna_version (default 0.6.0-pre when no embedded metadata pins it down), created_at (mtime of audit_log.json if present, otherwise current time), run_id (derived from the bundle directory name), workflow_hash (extracted from the first WorkflowStarted event in the audit log when readable; empty string when not), and components (sorted relative paths of every file in the bundle).
  • The synthesized object includes synthesized: true so downstream tooling can distinguish a reconstructed summary from a runner-native one.
  • When the bundle ALSO contains manifest.json (the canonical per-file checksum manifest from boruna-orchestrator), the migrator cross-checks it via verify_bundle and reports the result.

workflow-json

For a single workflow.json file:

  • If schema_version: 1 is present, the migrator reports a no-op and the file is byte-identical afterwards.
  • If schema_version is missing, the migrator adds it as 1.
  • If schema_version is present and greater than 1, the migrator errors out — downgrading from a future schema is out of scope for the beta.
  • If schema_version is some other value (negative, non-integer, string), the migrator rejects the input as malformed.

Note on field order: serde_json::Map in this workspace is a BTreeMap, so the migrated file’s keys land in lexicographic order. The semantic content is preserved; comments (which JSON does not support) are unaffected.

What the migrators do NOT do

  • Reconstruct missing checksums. A legacy evidence bundle that never had manifest.json cannot grow one with hashes that match the original artifacts after the fact — that would manufacture a false integrity guarantee. The synthesized bundle.json is a summary, not a checksum manifest.
  • Downgrade from a future schema. --to is forward-only in the beta.
  • Touch persistence stores (runs.db). Database schema migrations are deferred to a follow-up sprint that will land alongside the next breaking persistence change.
  • Edit examples/workflows/*/workflow.json retroactively. Those ship at the current schema with each release.

Coverage matrix (beta)

FromToevidence-bundleworkflow-json
0.5.0 (and earlier)0.6.0 / 1.0.0 / currentyes (best-effort summary)yes
0.6.0+ artifactscurrentno-opno-op
1.x+ artifactsdowngradenot supportednot supported

The matrix will expand in 1.x as additional breaking changes accumulate. Each future migrator will land with its own line in this table and a sprint reference in docs/roadmap.md.

Reporting issues

If boruna migrate rejects a bundle or workflow that you believe is valid 0.5.0 output, file an issue at https://github.com/escapeboy/boruna/issues with:

  • The full migrator output (run with --dry-run first).
  • A redacted copy of the bundle directory listing or workflow.json.
  • The Boruna version that produced the artifact.

Bundle Storage on S3

Sprint reference: post1-T-3.1.

The --bundle-storage s3://... flag tells boruna workflow run to copy the finalized evidence bundle to an S3 bucket after the local write succeeds. The local bundle remains the authoritative record; the S3 copy is an additional durable destination that survives the local data directory being wiped.

Status

  • Adapter shipped: S3 (this guide). Works against AWS S3 and S3-compatible endpoints (MinIO, Cloudflare R2, Backblaze B2, LocalStack).
  • Reserved schemes: gs:// (T-3.2, GCS) and azblob:// (T-3.3, Azure Blob) are rejected at parse time until those adapters ship.

Build with the s3 feature

The S3 adapter pulls in object_store + tokio + reqwest. To keep the default binary lean, the adapter is OFF by default. Enable it at build time:

# CLI binary (boruna)
cargo build --release --features boruna-cli/s3

# Direct orchestrator usage from another Rust crate
[dependencies]
boruna-orchestrator = { path = "...", features = ["s3"] }

When you build without the s3 feature and pass --bundle-storage s3://..., the URI rejects at parse time with a message that points you at the feature flag. Crucially: it never silently ignores the flag and ships an audit gap.

Configuring auth

object_store::aws::AmazonS3Builder::from_env() reads the standard AWS environment variables:

VariablePurpose
AWS_ACCESS_KEY_IDIAM access key
AWS_SECRET_ACCESS_KEYIAM secret
AWS_SESSION_TOKENSTS session token (optional)
AWS_REGIONDefaults to us-east-1 if unset
AWS_ENDPOINT_URLOverride for MinIO / R2 / LocalStack
AWS_ALLOW_HTTPSet to true for non-HTTPS endpoints (testing only)

For production AWS, prefer instance/role credentials (the SDK picks those up automatically when none of the env vars above are set).

Required IAM permissions

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "s3:PutObject",
        "s3:GetObject",
        "s3:ListBucket",
        "s3:DeleteObject"
      ],
      "Resource": [
        "arn:aws:s3:::your-bucket",
        "arn:aws:s3:::your-bucket/*"
      ]
    }
  ]
}

s3:DeleteObject is reserved for a future evidence prune flow; the current adapter does not delete objects.

Usage

Per-run

export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export AWS_REGION=eu-west-1

boruna workflow run examples/workflows/llm_code_review \
  --policy allow-all \
  --record \
  --bundle-storage s3://my-audit-bucket/prod/llm-review

After the run finalizes locally, the CLI prints:

evidence bundle: ./data/evidence/<run-id>
  bundle_hash: <hex>
  audit_log_hash: <hex>
  files: 12
  storage_ref: s3://my-audit-bucket/prod/llm-review/<run-id>

Via env var

The flag falls back to BORUNA_BUNDLE_STORAGE, so operators typically set this once in their service environment:

export BORUNA_BUNDLE_STORAGE=s3://my-audit-bucket/prod

URI shape

PatternEffect
s3://bucketObjects land at <run-id>/<file>
s3://bucket/prefixObjects land at prefix/<run-id>/<file>
s3://bucket/a/b/c/Trailing slash normalized; same as s3://bucket/a/b/c

The StorageRef returned by put is s3://bucket/prefix/<run-id>. Treat it as opaque; only the dispatcher parses it.

Reading bundles back

Once you have the storage_ref, use the orchestrator API to materialize the bundle into a local cache directory:

#![allow(unused)]
fn main() {
use boruna_orchestrator::audit::storage::{from_uri, StorageRef};

let storage = from_uri(Some("s3://my-audit-bucket/prod"))?.unwrap();
let local_dir = storage.get(&StorageRef("s3://my-audit-bucket/prod/<run-id>".into()))?;
// local_dir is now under BORUNA_BUNDLE_CACHE (defaults to <temp>/boruna-bundle-cache)
// and contains the same files the original run finalized locally.

boruna_orchestrator::audit::verify_bundle(&local_dir)?;
}

The cache directory is overwritten on each get. There is no automatic GC; sweep it periodically with find $BORUNA_BUNDLE_CACHE -mtime +N -delete or similar.

Failure semantics

ConditionBehavior
--bundle-storage URI invalidRun still completes; warning printed; storage_ref not recorded.
S3 put fails (network / auth / quota)Run still completes; warning printed; storage_ref not recorded. The local bundle is the authoritative record.
s3:// passed without s3 featureRun rejects at flag parse time with an actionable message.

A storage failure never masks a successful workflow. Conversely, operators who require S3 durability should monitor the warning output and re-run evidence verify against the local bundle to confirm it’s still on disk.

Error taxonomy

StorageError::Backend { kind, msg } uses these stable kinds for S3 operations:

kindMeaningRetry?
s3.transientNetwork blip, timeout, throttleYes — object_store already retries internally; bubbled up means retries exhausted.
s3.permanentAuth failure, NoSuchBucket, AccessDeniedNo — operator config issue.
s3.runtimeCould not build the tokio runtime backing the adapterNo — host issue.
s3.unexpected_keyObject listed under the run prefix but doesn’t match the expected path layoutNo — investigate; possible bucket pollution.

StorageError::NotFound(ref) fires when get is called against a ref that has zero objects under its prefix.

Testing against MinIO locally

# Spin up MinIO
docker run -p 9000:9000 -p 9001:9001 \
  -e MINIO_ROOT_USER=minioadmin \
  -e MINIO_ROOT_PASSWORD=minioadmin \
  quay.io/minio/minio server /data --console-address ":9001"

# Create a bucket via the console (http://localhost:9001) or mc
mc alias set local http://localhost:9000 minioadmin minioadmin
mc mb local/boruna-audit

# Run boruna against it
export AWS_ENDPOINT_URL=http://localhost:9000
export AWS_ACCESS_KEY_ID=minioadmin
export AWS_SECRET_ACCESS_KEY=minioadmin
export AWS_REGION=us-east-1
export AWS_ALLOW_HTTP=true

boruna workflow run examples/workflows/llm_code_review \
  --policy allow-all \
  --record \
  --bundle-storage s3://boruna-audit/local-test

The --features s3-it integration test under orchestrator/tests/ runs the full round-trip against a testcontainers-managed MinIO container; see orchestrator/tests/s3_integration.rs for the canonical example.

Determinism contract

storage_ref is operational metadata — it does not feed any audit-log hash or replay comparison. The bundle’s bundle_hash / audit_log_hash come from the local manifest and are independent of where the bundle is also stored.

Limitations

  • No automatic bucket creation. The bucket must exist before you point boruna at it. Failure to find the bucket surfaces as s3.permanent.
  • No multipart-upload tuning knob. object_store picks reasonable defaults (5 MB part size). Files smaller than that go via single PUT; larger ones use multipart automatically.
  • The cache directory is shared across adapters — if you rotate between two buckets that contain different bundles for the same run_id, the cache content is whichever was fetched most recently. Operators who care about cross-bucket consistency should set BORUNA_BUNDLE_CACHE to a per-bucket directory.
  • No retention / lifecycle policy is configured on the bucket; operators define this server-side via S3 lifecycle rules.

Bundle Storage on Google Cloud Storage

Sprint reference: post1-T-3.2.

The --bundle-storage gs://... flag tells boruna workflow run to copy the finalized evidence bundle to a GCS bucket after the local write succeeds. The local bundle remains the authoritative record; the GCS copy is an additional durable destination.

This adapter mirrors the S3 adapter — same trait, same BORUNA_BUNDLE_CACHE semantics, same failure contract, just with GCS auth + the gs:// scheme.

Status

  • Adapter shipped: GCS (this guide). Works against Google Cloud Storage and fsouza/fake-gcs-server for local testing.
  • Already shipped: S3 (T-3.1). See bundle-storage-s3.md.
  • Reserved scheme: azblob:// (T-3.3, Azure Blob) is rejected at parse time until that adapter ships.

Build with the gcs feature

# CLI binary (boruna)
cargo build --release --features boruna-cli/gcs

# Direct orchestrator usage from another Rust crate
[dependencies]
boruna-orchestrator = { path = "...", features = ["gcs"] }

# Combined with S3 if you operate multi-cloud
cargo build --release --features "boruna-cli/s3,boruna-cli/gcs"

When you build without the gcs feature and pass --bundle-storage gs://..., the URI rejects at parse time with a message that points you at the feature flag. Same UX guarantee as S3 — never silently ignored.

Configuring auth

object_store::gcp::GoogleCloudStorageBuilder::from_env() reads:

VariablePurpose
GOOGLE_SERVICE_ACCOUNT / GOOGLE_SERVICE_ACCOUNT_PATHPath to a JSON service-account key file
GOOGLE_SERVICE_ACCOUNT_KEYThe JSON service-account key inline (handy for K8s secrets)
GOOGLE_APPLICATION_CREDENTIALSApplication Default Credentials (Workload Identity, gcloud login)

In production, prefer Workload Identity (GKE) or Service-Account-Attached-to-VM (GCE) — neither requires a key on disk.

Required IAM permissions

Bind a service account with the predefined role roles/storage.objectAdmin on the bucket, or the more granular:

storage.buckets.get
storage.objects.create
storage.objects.delete
storage.objects.get
storage.objects.list

storage.objects.delete is reserved for a future evidence prune flow; the current adapter does not delete objects.

Usage

Per-run

export GOOGLE_APPLICATION_CREDENTIALS=/etc/boruna/sa.json
# Or for a key in env:
# export GOOGLE_SERVICE_ACCOUNT_KEY="$(cat /etc/boruna/sa.json)"

boruna workflow run examples/workflows/llm_code_review \
  --policy allow-all \
  --record \
  --bundle-storage gs://my-audit-bucket/prod/llm-review

After the run finalizes locally, the CLI prints:

evidence bundle: ./data/evidence/<run-id>
  bundle_hash: <hex>
  audit_log_hash: <hex>
  files: 12
  storage_ref: gs://my-audit-bucket/prod/llm-review/<run-id>

Via env var

export BORUNA_BUNDLE_STORAGE=gs://my-audit-bucket/prod

URI shape

PatternEffect
gs://bucketObjects land at <run-id>/<file>
gs://bucket/prefixObjects land at prefix/<run-id>/<file>
gs://bucket/a/b/c/Trailing slash normalized; same as gs://bucket/a/b/c

The StorageRef returned by put is gs://bucket/prefix/<run-id>. Treat it as opaque; only the dispatcher parses it.

Reading bundles back

#![allow(unused)]
fn main() {
use boruna_orchestrator::audit::storage::{from_uri, StorageRef};

let storage = from_uri(Some("gs://my-audit-bucket/prod"))?.unwrap();
let local_dir = storage.get(&StorageRef("gs://my-audit-bucket/prod/<run-id>".into()))?;
boruna_orchestrator::audit::verify_bundle(&local_dir)?;
}

The cache directory (default <temp>/boruna-bundle-cache, overridable via BORUNA_BUNDLE_CACHE) is shared with the S3 adapter — set per-bucket cache dirs if you operate multi-cloud and care about cross-bucket consistency.

Failure semantics

Same as S3 (see bundle-storage-s3.md). A storage failure never masks a successful workflow.

Error taxonomy

StorageError::Backend { kind, msg } uses these stable kinds for GCS operations:

kindMeaningRetry?
gcs.transientNetwork blip, timeout, throttleYes — object_store already retries internally; bubbled up means retries exhausted.
gcs.permanentAuth failure, NoSuchBucket, AccessDeniedNo — operator config issue.
gcs.runtimeCould not build the tokio runtime backing the adapterNo — host issue.
gcs.unexpected_keyObject listed under the run prefix but doesn’t match the expected path layoutNo — investigate; possible bucket pollution.

StorageError::NotFound(ref) fires when get is called against a ref that has zero objects under its prefix.

Testing against fake-gcs-server locally

# Spin up fake-gcs-server
docker run -p 4443:4443 \
  fsouza/fake-gcs-server:1.49.2 \
  -scheme http -host 0.0.0.0 -port 4443

# Create a bucket
curl -X POST 'http://localhost:4443/storage/v1/b?project=test-project' \
  -H 'Content-Type: application/json' \
  -d '{"name":"boruna-audit"}'

# Programmatic use from Rust:
use boruna_orchestrator::audit::storage_gcs::GcsBucketBuilder;
let store = GcsBucketBuilder::new("gs://boruna-audit/local-test")
    .with_endpoint("http://localhost:4443")
    .build()?;

(fake-gcs-server doesn’t need credentials; the adapter still calls from_env(), which simply doesn’t pick anything up — the endpoint override is what matters.)

The --features gcs-it integration tests under orchestrator/tests/ run the full round-trip against a testcontainers-managed fake-gcs-server container; see orchestrator/tests/gcs_integration.rs for the canonical example.

Determinism contract

storage_ref is operational metadata — it does not feed any audit-log hash or replay comparison. The bundle’s bundle_hash / audit_log_hash come from the local manifest and are independent of where the bundle is also stored.

Limitations

Same as S3: no automatic bucket creation, no multipart-upload tuning knob, shared cache directory across adapters, no retention/lifecycle policy (configure server-side via GCS Object Lifecycle rules).

Bundle Storage on Azure Blob Storage

Sprint reference: post1-T-3.3.

The --bundle-storage azblob://... flag tells boruna workflow run to copy the finalized evidence bundle to an Azure Blob container after the local write succeeds. The local bundle remains the authoritative record; the Azure copy is an additional durable destination.

This adapter mirrors the S3 and GCS adapters — same trait, same BORUNA_BUNDLE_CACHE semantics, same failure contract, just with Azure auth + the azblob:// scheme.

Status

  • Adapter shipped: Azure Blob Storage (this guide). Works against Azure proper and Azurite for local testing.
  • With T-3.3 landing, all three remote schemes (S3, GCS, Azure) ship. The BundleStorage trait can graduate from #[doc(hidden)] to pub in a future doc PR.

Build with the azure feature

# CLI binary (boruna)
cargo build --release --features boruna-cli/azure

# Direct orchestrator usage from another Rust crate
[dependencies]
boruna-orchestrator = { path = "...", features = ["azure"] }

# All three remote schemes together
cargo build --release --features "boruna-cli/s3,boruna-cli/gcs,boruna-cli/azure"

When you build without the azure feature and pass --bundle-storage azblob://..., the URI rejects at parse time with a message that points you at the feature flag. Same UX guarantee as S3 / GCS — never silently ignored.

Configuring auth

object_store::azure::MicrosoftAzureBuilder::from_env() reads:

VariablePurpose
AZURE_STORAGE_ACCOUNT_KEY / AZURE_STORAGE_ACCESS_KEYStorage account master key
AZURE_STORAGE_CLIENT_ID / AZURE_STORAGE_CLIENT_SECRET / AZURE_STORAGE_TENANT_IDService-principal OAuth
AZURE_STORAGE_SAS_KEYPre-shared SAS token

The account name comes from the URI (azblob://<account>/...), not the env, so a misconfigured AZURE_STORAGE_ACCOUNT_NAME won’t silently mismatch what the operator wrote.

In production, prefer Workload Identity (AKS) or Managed-Identity-Attached-to-VM — neither requires a key on disk.

Required RBAC roles

Bind a service principal or managed identity with the predefined role Storage Blob Data Contributor scoped to the container (or the storage account if you operate many containers per account).

URI shape

PatternEffect
azblob://account/containerObjects land at <run-id>/<file> inside container
azblob://account/container/prefixObjects land at prefix/<run-id>/<file>
azblob://account/container/a/b/c/Trailing slash normalized; same as .../a/b/c

Unlike S3 (single-namespace) and GCS (single-namespace), Azure has a two-level namespace: storage account + blob container. Both are encoded in the URI so an operator can grep their config and see exactly which account a bundle landed in.

The StorageRef returned by put is azblob://account/container/prefix/<run-id>. Treat it as opaque; only the dispatcher parses it.

Usage

Per-run

export AZURE_STORAGE_ACCOUNT_KEY="$(cat /etc/boruna/azure-key)"

boruna workflow run examples/workflows/llm_code_review \
  --policy allow-all \
  --record \
  --bundle-storage azblob://myacct/audit-bundles/prod

Via env var

export BORUNA_BUNDLE_STORAGE=azblob://myacct/audit-bundles/prod

Reading bundles back

#![allow(unused)]
fn main() {
use boruna_orchestrator::audit::storage::{from_uri, StorageRef};

let storage = from_uri(Some("azblob://myacct/audit-bundles/prod"))?.unwrap();
let local_dir = storage.get(&StorageRef(
    "azblob://myacct/audit-bundles/prod/<run-id>".into()
))?;
boruna_orchestrator::audit::verify_bundle(&local_dir)?;
}

The cache directory (default <temp>/boruna-bundle-cache, overridable via BORUNA_BUNDLE_CACHE) is shared with the S3 and GCS adapters — set per-bucket cache dirs if you operate multi-cloud and care about cross-bucket consistency.

Failure semantics

Same as S3/GCS (see bundle-storage-s3.md). A storage failure never masks a successful workflow.

Error taxonomy

StorageError::Backend { kind, msg } uses these stable kinds for Azure operations:

kindMeaningRetry?
azure.transientNetwork blip, timeout, throttleYes — object_store already retries internally; bubbled up means retries exhausted.
azure.permanentAuth failure, ContainerNotFound, AuthorizationFailureNo — operator config issue.
azure.runtimeCould not build the tokio runtime backing the adapterNo — host issue.
azure.unexpected_keyObject listed under the run prefix but doesn’t match the expected path layoutNo — investigate; possible container pollution.

StorageError::NotFound(ref) fires when get is called against a ref that has zero objects under its prefix.

Testing against Azurite locally

Programmatic use from Rust against an Azurite container:

#![allow(unused)]
fn main() {
use boruna_orchestrator::audit::storage_azure::AzureBlobBucketBuilder;

let store = AzureBlobBucketBuilder::new(
    "azblob://devstoreaccount1/boruna-audit/local-test"
)
.with_use_emulator(true)
.with_endpoint("http://localhost:10000/devstoreaccount1")
.build()?;
}

with_use_emulator(true) switches the SDK into Azurite mode (well-known account devstoreaccount1, well-known shared key). with_endpoint lets the SDK reach a non-default port.

Container creation against Azurite needs a SharedKey-signed PUT to ?restype=container — out of scope for the bundled adapter, which assumes the container exists. In a real test, create the container with the Azure CLI (az storage container create) or the Azure Storage Explorer.

Determinism contract

storage_ref is operational metadata — it does not feed any audit-log hash or replay comparison. The bundle’s bundle_hash / audit_log_hash come from the local manifest and are independent of where the bundle is also stored.

Limitations

  • No automatic container creation. The container must exist before you point boruna at it. Failure to find it surfaces as azure.permanent. Create with az storage container create --name <container> --account-name <account>.
  • No integration test in this build. Azurite requires SharedKey-signed container-create calls and object_store’s Azure adapter doesn’t expose a create_container primitive. Adding it would require either pulling in the full azure-storage crate or hand-rolling the SharedKey signing for a single test. The 18 unit tests cover URI parsing, path concatenation, ref-to-run-id extraction, and error classification — the Boruna-side logic. Verification against real Azure is the operator follow-up.
  • No multipart-upload tuning knob. object_store picks reasonable defaults.
  • Shared cache directory across adapters; set per-bucket if you care.
  • No retention/lifecycle policy is configured; use Azure Storage’s built-in lifecycle management.

KEK Rotation

Sprint reference: post1-T-2.4.

boruna evidence rotate-kek re-wraps the data-encryption-key (DEK) of one or more encrypted evidence bundles under a new key-encryption-key (KEK). The DEK itself is not changed — only the manifest’s wrapped_dek, wrapped_dek_nonce, and kek_id fields. Every per-file AES-GCM authentication tag in the bundle remains valid because the keying material that produced those tags hasn’t moved.

When to rotate

  • A KEK is suspected leaked or has been retired in your KMS.
  • A scheduled rotation policy (e.g. “rotate every 6 months”).
  • Migrating bundles from a legacy kek_id to a new one.

Compliance auditors typically only need to see that rotation happened — the bundle’s bundle_hash updates because the manifest contents change, but the audit log inside the bundle is unchanged.

Single-bundle rotation

boruna evidence rotate-kek \
  /path/to/<run-id> \
  --old-kek $(printenv OLD_KEK_HEX) \
  --new-kek $(printenv NEW_KEK_HEX) \
  --kek-id-to "rotation-2026-q3"

The bundle’s manifest.json is rewritten atomically (sibling tmp + rename). The rest of the bundle directory is untouched.

Batch rotation

Point --target at a parent directory whose immediate subdirectories are bundles:

boruna evidence rotate-kek \
  /var/lib/boruna/evidence \
  --old-kek $OLD --new-kek $NEW \
  --kek-id-to "rotation-2026-q3" \
  --parallelism 4

Each bundle is processed in its own rayon task, bounded by --parallelism (default min(8, num_cpus)). Per-bundle failures do not abort the batch — already-rotated bundles stay rotated, and the failed bundle is reported on stderr. The CLI exits non-zero if any bundle failed.

Dry-run

Always dry-run first on a representative bundle:

boruna evidence rotate-kek bundles/abc... \
  --old-kek $OLD --new-kek $NEW --kek-id-to new-id \
  --dry-run

Dry-run validates that the old KEK can unwrap, that the new KEK re-wraps cleanly, and prints the planned kek_id change. No files are modified.

kek_id_from filter

Use --kek-id-from <id> to defend against accidental double-rotation in a mixed-state batch (some bundles already rotated, others not):

boruna evidence rotate-kek bundles/ \
  --old-kek $OLD --new-kek $NEW \
  --kek-id-from "rotation-2026-q2" \
  --kek-id-to "rotation-2026-q3"

Bundles whose current kek_id does not match --kek-id-from are reported as failures (bundle kek_id is 'X'; --kek-id-from is 'Y') without being modified.

Verifying after rotation

boruna evidence verify /path/to/<run-id> \
  --bundle-encryption-key $NEW_KEK_HEX

Verifying with the old KEK after rotation MUST fail with evidence.encryption_key_mismatch — that’s the regression the rotation tool’s tests assert on.

Spec reference

The 1.0 evidence-bundle reader contract already accommodates re-wrapping: per-file ciphertext stays valid because the DEK is unchanged. See docs/spec/evidence-bundle-1.0.md. No spec version bump is required for this feature.

CLI Reference

The boruna binary is the primary interface for compiling, running, inspecting, and managing .ax programs and workflows.

Version

boruna --version (or -V) prints boruna <version>. The other binaries (boruna-mcp, boruna-pkg, boruna-orch) do the same.

Installation

cargo build --workspace
# Binary at: target/debug/boruna

To use boruna without cargo run --bin boruna --, add target/debug/ to your PATH or install with:

cargo install --path crates/llmvm-cli

Top-level commands

boruna <command> [options]

Commands:
  compile     Compile a .ax file to bytecode
  run         Run a .ax file
  trace       Run and emit an execution trace
  replay      Replay from a recorded event log
  inspect     Inspect a compiled module
  ast         Print the AST for a .ax file
  lang        Language diagnostics, repair, and the code registry
  doctor      Environment and toolchain health checks
  size        Bytecode artifact size report for a .ax file
  framework   Framework app validation and testing
  workflow    Workflow validation, execution, and graph inspection
  evidence    Evidence bundle inspection and verification
  template    Template listing and application
  skills      Embedded, agent-curated documentation
  trace2tests Generate regression tests from traces

boruna compile

Compile a .ax source file to bytecode.

boruna compile <file.ax>

Outputs the compiled module summary (functions, capabilities declared). Does not execute.


boruna run

Compile and run a .ax file.

boruna run <file.ax> [options]

Options:
  --policy <name>    Capability policy: allow-all, deny-all (default: deny-all)
  --record           Write an event log to .boruna/runs/<id>/
  --live             Enable real capability handlers (requires http feature)
  --trace            Emit a full execution trace to stdout
  --step-limit <n>   Abort if execution exceeds n steps
  --watch            Re-run on every change to the file (post-1.0)

Examples:

# Run in demo mode (no capabilities)
boruna run examples/hello.ax

# Run with all capabilities allowed
boruna run app.ax --policy allow-all

# Run with real HTTP (requires --features boruna-cli/http)
cargo run --features boruna-cli/http --bin boruna -- run app.ax --policy allow-all --live

# Watch mode — re-run on every save until Ctrl-C
boruna run app.ax --watch

Watch mode

--watch re-executes the file on every change. Filesystem events are debounced to 200ms, so a single editor save triggers exactly one rerun even on platforms that emit a flurry of events per save.

A separator line marks each rerun:

── reloading app.ax at 14:03:11 ──
3
steps: 4

A failed run (compile error or runtime panic) does not exit watch mode — the error is printed and the watcher waits for the next save. Press Ctrl-C to exit.

Watch combines with the standard run flags: boruna run app.ax --watch --policy ./policy.json --record events.json.


boruna trace

Run a .ax file and emit a step-by-step execution trace.

boruna trace <file.ax> [--policy <name>]

boruna replay

Replay execution from a recorded event log.

boruna replay <file.ax> --from <event-log-path> [--verify]

Options:
  --verify    Compare replay output to recorded output; fail if they differ

boruna inspect

Inspect a compiled bytecode module.

boruna inspect <file.ax>

Prints: function table, constant pool, declared capabilities, bytecode disassembly.


boruna ast

Print the parsed AST for a .ax file as JSON.

boruna ast <file.ax> [--json]

boruna lang

Language diagnostics and auto-repair.

boruna lang check <file.ax> [--json]
boruna lang repair <file.ax>
boruna lang codes [--json]

Subcommands:
  check     Run diagnostics: type errors, undeclared capabilities, unreachable code
  repair    Apply auto-repair suggestions from diagnostics
  codes     List the registry of stable diagnostic codes (E001–E010)

Examples:

# Check for errors with machine-readable output
boruna lang check app.ax --json

# Automatically repair issues
boruna lang repair app.ax

# Resolve a diagnostic code seen in `lang check --json` output
boruna lang codes --json

lang codes emits the registry from docs/reference/diagnostic-codes.md. Codes are stable forever — tools and agents may switch on them.


boruna doctor

Environment and toolchain health checks.

boruna doctor [--json]

Reports the binary version, which optional features were compiled in, whether a Rust toolchain is reachable, the persistent data directory’s writability, and whether the current directory looks like a Boruna project root. Read-only. Exits 1 if any check has error status.


boruna size

Bytecode artifact size report for a .ax source file.

boruna size <file.ax> [--json]

Compiles the file and reports per-function opcode counts, module-wide totals (functions, ops, constants, types, globals), and the serialized .axbc artifact byte size. Nothing is written to disk.


boruna confidence

Calibrated confidence for approval gates. See docs/design-conformal-gating.md.

boruna confidence threshold <calibration.json> --alpha-permille <1..999> [--score <0..1000>] [--json]

Reads a labelled calibration set and prints the score at or above which a gate may skip a human, so that a wrong answer is auto-approved with probability at most alpha. Also prints how the calibration set itself would have behaved at that threshold, and the decision for --score when given. When the set has too few wrong examples to certify alpha (at least ceil(1000 / alpha_permille) - 1 are needed), the threshold is never and every case goes to a human. Exits 1 on an invalid file or alpha.


boruna skills

Embedded, agent-curated documentation — compiled into the binary so an agent can learn Boruna from the installed binary alone.

boruna skills list [--json]
boruna skills get <name> [--json]
boruna skills emit <dir>
boruna skills pack "<query>" [--budget <tokens>] [--json]

Subcommands:
  list   List available skill documents
  get    Print one skill document (ax-language, cli, workflows, diagnostics)
  emit   Write every skill as <dir>/<name>/SKILL.md (name + description frontmatter)
  pack   Return only the sections relevant to a query, within a token budget

skills get exits 1 on an unknown skill name and lists the available names.

The cli skill ends with a command reference generated from the installed binary, so it lists exactly the commands and flags that exist. A unit test fails when a hand-written skill names a boruna <command> that does not exist.

skills emit is idempotent and overwrites existing files. skills pack ranks sections by keyword overlap with the query (ties broken by skill name and position), so the same query always returns the same bytes. --budget is an estimate (characters / 4, default 2000); the best section is truncated rather than dropped if it alone exceeds it. Exits 1 when nothing matches.


boruna framework

Validate and test framework apps (Elm-architecture .ax apps with init/update/view).

boruna framework validate <file.ax>
boruna framework test <file.ax> [options]

Options for test:
  -m <messages>    Comma-separated message sequence, e.g. "increment:1,reset:0"

Examples:

boruna framework validate examples/framework/counter_app.ax
boruna framework test examples/framework/counter_app.ax -m "increment:1,increment:1,reset:0"

boruna workflow

Validate and run workflow DAGs.

boruna workflow validate <workflow-dir/>
boruna workflow run <workflow-dir/> [options]

Options for run:
  --policy <name>    Capability policy (default: deny-all)
  --record           Write evidence bundle to .boruna/runs/<id>/
  --live             Enable real capability handlers
  --replay <dir>     Replay from an existing evidence bundle
  --verify           (with --replay) Verify outputs match recorded values

Examples:

# Validate DAG structure
boruna workflow validate examples/workflows/llm_code_review

# Run in demo mode
boruna workflow run examples/workflows/llm_code_review --policy allow-all

# Run and record evidence
boruna workflow run examples/workflows/llm_code_review --policy allow-all --record

# Run with real HTTP and LLM calls
cargo run --features boruna-cli/http --bin boruna -- \
  workflow run examples/workflows/llm_code_review \
  --policy allow-all --live --record

boruna workflow schedule

Run a workflow on a cron schedule. Loops until SIGINT (Ctrl-C).

boruna workflow schedule <workflow-dir/> [options]

Options:
  --cron <expr>            5-field cron expression (required), e.g. "*/5 * * * *"
  --policy <name>          Capability policy (default: deny-all)
  --data-dir <dir>         Directory for runs.db and per-run output (default: .boruna/data)
  --max-concurrency <n>    Maximum concurrent runs; skips tick if a run is already active (default: 1)
  --live                   Enable real capability handlers (requires http feature)

When a scheduled tick fires while a previous run is still active, that tick is skipped.

boruna workflow eval

Run a workflow against two LLM provider configurations and compare evidence bundles.

boruna workflow eval <workflow-dir/> --providers-a <file.json> --providers-b <file.json> [options]

Options:
  --providers-a <file>     First provider config JSON file (required)
  --providers-b <file>     Second provider config JSON file (required)
  --runs <n>               Runs per provider (default: 1)
  --data-dir <dir>         Directory for evidence bundles
  --json                   Machine-readable output

Reports per-step output agreement and timing differences between the two provider sets.

boruna workflow find

Recursively discover workflow.json files under a directory tree.

boruna workflow find [dir] [--json]

Validates each discovered workflow and prints path, name, step count, and validity. dir defaults to the current directory. Use --json for a JSON array.


boruna workflow graph

Emit the workflow DAG as structured graph facts.

boruna workflow graph <dir> [--json]

Reports nodes (each step’s kind, capabilities, and dependencies), edges, topological execution order, roots (steps with no dependencies), and leaves (steps nothing depends on). Read-only — only workflow.json is read, step source files are not. Exits 1 if the graph contains a cycle (is_dag: false).


boruna evidence

Inspect, verify, and manage evidence bundles.

boruna evidence create <run-id> --output-dir <dir> [--data-dir <dir>]
boruna evidence verify <bundle-dir/> [--bundle-encryption-key <hex>]
boruna evidence inspect <bundle-dir/> [--json] [--decrypt] [--bundle-encryption-key <hex>]
boruna evidence diff <bundle-a> <bundle-b> [--json]
boruna evidence gc-blobs [--data-dir <dir>] [--dry-run] [--json]
boruna evidence rotate-kek <target> --old-kek <hex> --new-kek <hex> [options]

Examples:

# Build a bundle from a completed run
boruna evidence create abc123def456 --output-dir ./bundles

# Inspect a bundle
boruna evidence inspect .boruna/runs/20260315-143022-abc4d/
boruna evidence inspect .boruna/runs/20260315-143022-abc4d/ --json

# Verify a bundle
boruna evidence verify .boruna/runs/20260315-143022-abc4d/

evidence inspect

For plaintext (non-encrypted) bundles, evidence inspect automatically shows step output content from outputs/<step_id>/result.json. Each step output is truncated at 500 characters in text mode. With --json, the full parsed content appears under the "step_outputs" key. Encrypted bundles print a hint to stderr if --decrypt / --bundle-encryption-key is not supplied.

evidence create

Build an evidence bundle from a persisted run. Reads the run’s metadata, step checkpoints, and hash-chained audit log; writes a bundle directory with workflow.json, policy.json, per-step outputs, audit_log.json, env_fingerprint.json, and manifest.json. Bundles are created on demand — the runner does not auto-create them.

boruna evidence create <run-id> --output-dir <dir> [--data-dir <dir>]

The bundle is written to <output-dir>/<run-id>/.

evidence diff

Compare two evidence bundles side-by-side.

boruna evidence diff <bundle-a> <bundle-b> [--json]

Reports differences in workflow metadata, step outputs, audit event counts, and verification status. Use --json for machine-readable output.

Examples:

boruna evidence diff .boruna/runs/run-baseline/ .boruna/runs/run-rerun/
boruna evidence diff baseline/ rerun/ --json

evidence gc-blobs

Sweep orphaned content-addressed blobs from the data directory.

boruna evidence gc-blobs [--data-dir <dir>] [--dry-run] [--json]

An orphan is a blob file no longer referenced by any run checkpoint. Reports {deleted, skipped, bytes_freed}. Use --dry-run to report without deleting.

evidence rotate-kek

Rotate the key-encryption key (KEK) on one or more encrypted bundles without re-encrypting file content.

boruna evidence rotate-kek <target> --old-kek <hex> --new-kek <hex> [options]

Options:
  --kek-id-from <id>     Only rotate bundles whose current kek_id matches (safety check)
  --kek-id-to <id>       kek_id written to the rotated manifest (default: "default")
  --dry-run              Print planned actions without modifying any bundle
  --parallelism <n>      Parallel bundle limit in batch mode (default: min(8, num_cpus))

<target> may be a single bundle directory or a parent directory whose immediate subdirectories are bundles (batch mode).


boruna template

List and apply app templates.

boruna template list
boruna template apply <name> [options]

Options for apply:
  --args <key=value,...>    Template variable substitutions
  --validate                Validate the generated output after applying

Examples:

boruna template list
boruna template apply crud-admin --args "entity_name=products,fields=name|price" --validate

Available templates: crud-admin, form-basic, auth-app, realtime-feed, offline-sync


boruna trace2tests

Generate regression tests from execution traces.

boruna trace2tests <trace-file> --output <test-dir/>

See TRACE_TO_TESTS.md for details.


Global options

  --help      Print help for any command
  --version   Print the Boruna version

Command reference

Generated from boruna 3.5.0 --help when this site was built. For explanations and examples see the CLI guide.

boruna compile

  • boruna compile <FILE> [--output] — Compile a .ax source file to bytecode

boruna run

  • boruna run <FILE> [--policy] [--max-steps] [--record] [--live] [--record-net-to] [--replay-net-from] [--watch] [--providers] — Run a .ax source file or bytecode file

boruna trace

  • boruna trace <FILE> — Run with execution tracing enabled

boruna replay

  • boruna replay <FILE> <LOG> — Replay execution from a recorded event log

boruna inspect

  • boruna inspect <FILE> — Inspect a bytecode file

boruna ast

  • boruna ast <FILE> — Dump the AST of a .ax source file

boruna fmt

  • boruna fmt <FILE> [--check] — Format a .ax source file (canonical pretty-print)

boruna framework

  • boruna framework — Framework commands
  • boruna framework new <NAME> [--dir] — Create a new framework app from template
  • boruna framework validate <FILE> — Validate a .ax file conforms to the App protocol
  • boruna framework test <FILE> [--messages] — Run a framework app interactively with messages
  • boruna framework inspect-state <FILE> [--messages] — Inspect framework app state after running messages
  • boruna framework simulate <FILE> <MESSAGES> — Simulate a sequence of messages and display state transitions
  • boruna framework inspect <FILE> [--json] — Print App contract summary (State, Messages, Effects) — machine-readable
  • boruna framework diag <FILE> [--messages] — Structured diagnostics output (JSON)
  • boruna framework trace-hash <FILE> [--messages] — Run with tracing and print a stable hash of the trace
  • boruna framework replay <FILE> <LOG> — Replay a recorded cycle log and verify determinism

boruna lang

  • boruna lang — Language tooling commands (diagnostics, repair)
  • boruna lang check <FILE> [--json] [--output] — Check a source file and report diagnostics
  • boruna lang repair <FILE> [--from] [--apply] — Repair a source file using diagnostic suggestions
  • boruna lang codes [--json] — List the registry of stable diagnostic codes
  • boruna lang caps <FILE> [--json] — Report each function’s declared vs. inferred-needed capabilities and flag over-declarations (capabilities granted but never used)

boruna doctor

  • boruna doctor [--json] — Environment and toolchain health checks

boruna size

  • boruna size <FILE> [--json] — Report the bytecode artifact size of a .ax source file

boruna skills

  • boruna skills — Embedded, agent-curated documentation (list, get, emit, pack)
  • boruna skills list [--json] — List available agent skill documents
  • boruna skills get <NAME> [--json] — Print an agent skill document by name
  • boruna skills emit <DIR> — Write every skill as <DIR>/<name>/SKILL.md for an agent to load
  • boruna skills pack <QUERY> [--budget] [--json] — Return only the skill sections relevant to a query, within a token budget

boruna confidence

  • boruna confidence — Calibrated confidence for approval gates (threshold)
  • boruna confidence threshold <FILE> [--alpha-permille] [--score] [--json] — Show the auto-approve threshold a calibration file gives for a target false approval rate, and optionally the decision for one score

boruna trace2tests

  • boruna trace2tests — Trace-to-test tools (record, generate, run, minimize)
  • boruna trace2tests record <FILE> [--messages] [--out] — Record an execution trace from a framework app
  • boruna trace2tests generate [--trace] [--name] [--out] — Generate a test spec from a recorded trace
  • boruna trace2tests run [--spec] [--source] — Run a test spec against source code
  • boruna trace2tests minimize [--trace] [--source] [--predicate] [--out] — Minimize a failing trace using delta debugging

boruna template

  • boruna template — Template tools (list, apply, validate)
  • boruna template list [--dir] — List available templates
  • boruna template apply <NAME> [--dir] [--args] [--out] [--validate] — Apply a template with arguments

boruna literate

  • boruna literate — Literate workflow specs — extract embedded .ax/.qnt code fences from a markdown narrative into per-file outputs. See docs/architecture-literate-workflows.md
  • boruna literate extract <FILE> [--out-dir] [--json] [--verbose] — Extract <lang> <filename> += code fences from a markdown document into per-file outputs. Accepted languages: ax, boruna, quint. Other fences (rust, bash, …) are ignored. See docs/architecture-literate-workflows.md

boruna repl

  • boruna repl <FILE> [--policy] — Interactive REPL for .ax modules — load a file, evaluate expressions interactively, inspect the environment. See docs/architecture-boruna-repl.md

boruna simulate

  • boruna simulate <DIR> [--max-samples] [--seed] [--policy] [--invariant] [--witnesses] [--json] — Random property-based simulation of a workflow. Runs the workflow N times under a user-supplied invariant (and optional witnesses) and reports violation count + witness frequencies. See docs/architecture-boruna-simulate.md

boruna new

  • boruna new <TEMPLATE> [--dir] [--target] [--var] [--no-input] [--force] — Scaffold a new project from a template (interactive)

boruna workflow

  • boruna workflow — Workflow execution and validation
  • boruna workflow validate <DIR> [--print-hash] — Validate a workflow definition directory
  • boruna workflow run <DIR> [--policy] [--record] [--evidence-dir] [--encrypt-bundle] [--bundle-encryption-key] [--bundle-kek-id] [--live] [--data-dir] [--ephemeral] [--concurrency] [--skip-if-running] [--submit-only] [--expect-workflow-hash] [--bundle-storage] [--providers] — Run a workflow
  • boruna workflow approve <RUN_ID> <STEP_ID> [--data-dir] — Approve a paused approval-gate step. Records an approval sentinel in the run’s metadata; the operator must run boruna workflow resume <run-id> afterward to advance the run past the gate
  • boruna workflow reject <RUN_ID> <STEP_ID> [--reason] [--data-dir] — Reject a paused approval-gate step. Records a rejection sentinel; boruna workflow resume <run-id> will then halt the run as Failed with the optional reason as the error message
  • boruna workflow trigger <RUN_ID> <STEP_ID> [--token] [--payload] [--payload-file] [--data-dir] — Trigger a paused external_trigger step (sprint 0.3-S15). Records the supplied payload as the step’s output and primes resume to advance past the gate. Operator must run boruna workflow resume <run-id> afterward to actually execute downstream steps
  • boruna workflow show <RUN_ID> [--json] [--data-dir] — Show the full state of a single run: row, step checkpoints, and approval-gate decisions. Use --json for machine-readable output (jq-friendly). Reads from the same --data-dir as run/resume
  • boruna workflow list [--status] [--json] [--data-dir] — List runs in the persistent store. Optional –status filter
  • boruna workflow resume <RUN_ID> [--data-dir] [--workflow-dir] [--policy] [--live] [--concurrency] [--expect-workflow-hash] — Resume a previously-paused or crashed workflow run by id
  • boruna workflow schedule <DIR> [--cron] [--policy] [--data-dir] [--max-concurrency] [--live] — Run a workflow on a cron schedule in a long-running daemon process. Validates the cron expression on startup (fail fast), then loops: sleep until next fire time → invoke the runner API → log outcome. Ctrl-C / SIGTERM finish any in-progress run then exit cleanly
  • boruna workflow eval <WORKFLOW_DIR> [--providers-a] [--providers-b] [--runs] [--data-dir] [--json] — Run the same workflow against two LLM provider configs and compare outputs
  • boruna workflow find <DIR> [--json] — Find and inspect workflow definitions under a directory tree
  • boruna workflow graph <DIR> [--json] — Emit the workflow DAG as structured graph facts

boruna evidence

  • boruna evidence — Evidence bundle inspection and verification
  • boruna evidence create <RUN_ID> [--output-dir] [--data-dir] — Build an evidence bundle from a persisted run (sprint 0.4-S10). Reads the run’s metadata, step checkpoints, and hash-chained audit log; writes a bundle directory containing workflow.json, policy.json, per-step outputs, audit_log.json, env_fingerprint.json, and a manifest.json with bundle hash + per-file checksums
  • boruna evidence verify <DIR> [--bundle-encryption-key] [--expected-bundle-hash] [--require-encryption] [--verify-key] [--require-signature] — Verify an evidence bundle for integrity
  • boruna evidence inspect <DIR> [--json] [--itf] [--decrypt] [--bundle-encryption-key] — Inspect an evidence bundle’s manifest
  • boruna evidence gc-blobs [--data-dir] [--dry-run] [--json] — Sweep orphan blobs from the data-dir’s blobs/ tree (sprint W3-B). An orphan is a content-addressed blob file no longer referenced by any step_checkpoints.output_blob_ref row. Reports {deleted, skipped, bytes_freed}. Holds an exclusive write lock on runs.db for the duration of the sweep — see docs/design-blob-gc.md for the TOCTOU rationale
  • boruna evidence rotate-kek <TARGET> [--old-kek] [--new-kek] [--kek-id-from] [--kek-id-to] [--dry-run] [--parallelism] — Rotate the KEK on one or more encrypted evidence bundles (post1-T-2.4). Unwraps the DEK with the old KEK, re-wraps it under the new KEK, and atomically rewrites each bundle’s manifest.json. Per-file ciphertext is unchanged because the DEK itself does not change
  • boruna evidence redact <DIR> [--event] [--field] [--reason] — Verifiably redact one audit-log entry in an evidence bundle so PII can be removed from a SEALED bundle without breaking verification
  • boruna evidence diff <BUNDLE_A> <BUNDLE_B> [--json] — Compare two evidence bundles side-by-side (post1-evidence-diff). Reports differences in workflow metadata, step outputs, audit event counts, and verification status
  • boruna evidence attest <DIR> [--verify] [--signing-key] [--verify-key] [--output] — Emit (or verify) an in-toto Statement + DSSE envelope for the bundle’s runtime provenance, for interop with the supply-chain ecosystem (cosign verify-blob, in-toto-verify). Additive — does NOT touch the native bundle format. Writes attestation.intoto.dsse.json into the bundle directory
  • boruna evidence anchor <DIR> [--rekor-url] [--offline] [--verify] [--output] — Anchor a signed bundle in a Sigstore Rekor transparency log, adding an external witness + trusted timestamp on top of the bundle’s own hash chain (closes the “trust the recorder / silent backdating” hole). Builds a hashedrekord entry from the manifest’s bundle_hash + ed25519 signature. Three modes:
  • boruna evidence report <DIR> [--framework] [--format] — Generate a human-readable COMPLIANCE evidence-mapping report that maps a bundle’s actual contents to the specific regulatory obligation each one helps satisfy. Verifies the bundle first and stamps the verdict at the top; a tampered/unverifiable bundle produces a report that says so loudly. This is a technical mapping, NOT a certificate of compliance
  • boruna evidence otel <DIR> [--out] — Export the bundle’s execution as OpenTelemetry spans in OTLP/JSON — the file format any OTel collector ingests. No SDK dependency, no network: emit the document and POST it to a collector (or pipe it through the otlpjson file receiver) to surface the run in Jaeger, Tempo, Honeycomb, Datadog, etc

boruna capability

  • boruna capability — Capability surface inspection (versioned identity for caching)
  • boruna capability list [--json] — List all capabilities this binary exposes, with stable identity hash. Use capability_set_hash as part of a cache key to safely memoize deterministic results across binary upgrades. See docs/reference/capability-identity.md

boruna metrics

  • boruna metrics — Prometheus metrics export from the persistent run store (sprint 0.4-S12). See docs/design-prometheus-metrics.md for the architectural decision and operator integration pattern (cron + node_exporter’s textfile collector)
  • boruna metrics export [--data-dir] — Export current metrics in Prometheus text format to stdout. Pipe to a .prom file under node_exporter’s textfile collector directory:

boruna policy

  • boruna policy — Policy file validation and inspection (sprint 0.4-S15). See docs/design-policy-as-code.md and docs/reference/policy-schema.md for the schema and the stable error_kind taxonomy
  • boruna policy validate <FILE> [--json] — Strict-validate a policy file. Exits 0 on ok, 2 on validation error, 1 on file IO error. Designed as a CI gate
  • boruna policy show <FILE> — Validate then print the effective policy in human-readable form: default behavior, denormalized rule list, net policy bounds

boruna migrate

  • boruna migrate <KIND> <PATH> [--from] [--to] [--dry-run] [--in-place] — Migration tooling beta (sprint W5-C). Upgrades pre-1.0 Boruna artifacts to the current on-disk format. See docs/guides/migration.md for the coverage matrix and recommended workflow

.ax Language Reference

Looking for the formal specification? This page is the narrative, example-driven reference. The authoritative grammar, type rules, and capability semantics live in docs/spec/ax-language-1.0.md. When the two disagree, the spec wins.

.ax is Boruna’s statically-typed, deterministic scripting language. It compiles to Boruna bytecode and runs on the Boruna VM. It is designed for workflow steps: small, focused, pure functions with explicit capability declarations for any side effects.

File structure

Every standalone .ax file must define fn main() -> Int. The main function is the entry point for boruna run.

fn main() -> Int {
    42
}

Types

TypeDescriptionExample
Int64-bit signed integer42, -7
Float64-bit float3.14
StringUTF-8 string"hello"
BoolBooleantrue, false
UnitNo value()
Option<T>Optional valueSome(42), None
Result<T, E>Success or errorOk(42), Err("msg")
List<T>Ordered list[1, 2, 3]
Map<K, V>Key-value map{"a": 1, "b": 2}
RecordsNamed fieldsPoint { x: 1, y: 2 }
EnumsTagged unionShape::Circle(5)

Variables

Use let with an explicit type annotation:

let name: String = "Boruna"
let count: Int = 0
let flag: Bool = true

A binding that you want to change later is declared with let mut and rebound with =:

let mut total: Int = 0
total = total + 5

Rebinding changes what the name refers to; values themselves (records, lists, maps) are never modified in place. Reassigning a binding declared without mut still compiles, but boruna lang check reports warning E010 and boruna lang repair adds the missing mut. It will be a compile error in language version 2.0.

Loops

fn sum(items: List<Int>) -> Int {
    let mut total: Int = 0
    for x in items {
        total = total + x
    }
    total
}

fn factorial(n: Int) -> Int {
    let mut result: Int = 1
    let mut i: Int = n
    while i > 0 {
        result = result * i
        i = i - 1
    }
    result
}

for iterates a List in order; the loop variable and any let inside the body are scoped to the body. A loop that never ends is stopped by the step limit (--step-limit) with a runtime error. Recursion still works and is often the clearer choice.

No semicolons. Each statement is on its own line.

Functions

fn add(a: Int, b: Int) -> Int {
    a + b
}

The last expression in a function body is the return value. No return keyword needed.

Capability annotations

Functions that perform side effects must declare the required capabilities:

fn fetch(url: String) -> String !{net.fetch} {
    // live implementation
}

fn call_model(prompt: String) -> String !{llm.call} {
    // live implementation
}

// Multiple capabilities
fn fetch_and_cache(url: String) -> String !{net.fetch, fs.write} {
    // live implementation
}

Without the annotation, the VM will reject any attempt to call the capability at runtime.

Intent declarations

A function may declare a single machine-read purpose with an intent "..." clause after its signature. The clause is optional and order-independent with capability and contract clauses:

fn transfer(amount: Int) -> Int !{db.write} intent "Move funds between accounts" {
    // implementation
}

Intent is captured into the run’s evidence bundle (intents.json, keyed by step id) so an auditor sees what each step was authorized to do alongside what it actually did. It is covered by the bundle checksums, so tampering with a captured intent makes boruna evidence verify fail. A function may declare at most one intent; a second is a compile error.

Contracts (requires)

A function may declare one or more requires <expr> preconditions, checked at runtime against its arguments on entry:

fn transfer(amount: Int) -> Int !{db.write} requires amount > 0 {
    // runs only if amount > 0
}

If a precondition is false when the function is called, execution traps with a contract violation carrying a counterexample — the concrete arguments that triggered it (e.g. [0]) — so the failing input is reproducible. In a workflow, the violation surfaces with the stable error_kind contract_violation and the counterexample is recorded in the run’s hash-chained audit log (tamper-evident evidence). A violation is deterministic in the inputs, so it is not retry-eligible.

Contracts are enforced purely at runtime (concrete-trace checking) — Boruna does not use SMT/symbolic proving. ensures postconditions are parsed but not yet enforced.

Records

Define named record types with the type keyword:

type Point {
    x: Int,
    y: Int,
}

let p: Point = Point { x: 3, y: 4 }
let px: Int = p.x

Record spread creates an updated copy:

let p2: Point = Point { ..p, y: 10 }

Enums

Each variant is either a unit variant or carries a single payload value. Construct a value with the EnumName::Variant(payload) form (unit variants take no parentheses):

enum Shape {
    Circle(Float),
    Square(Float),
}

let s: Shape = Shape::Circle(5.0)

Pattern matching

let result: String = match s {
    Circle(radius) => "circle"
    Square(side) => "square"
    _ => "unknown"
}

Match on Option:

let value: Option<Int> = Some(42)
let n: Int = match value {
    Some(x) => x
    None => 0
}

Match on Result:

let r: Result<Int, String> = Ok(99)
let out: Int = match r {
    Ok(v) => v
    Err(_) => -1
}

Match on strings:

let greeting: String = match lang {
    "en" => "hello"
    "es" => "hola"
    _ => "hi"
}

Conditionals

let label: String = if score > 90 {
    "pass"
} else {
    "fail"
}

Lists

let items: List<Int> = [1, 2, 3, 4, 5]

List operations are available through the standard library.

Maps

let config: Map<String, Int> = { "timeout": 30, "retries": 3 }

Framework apps

Framework apps implement the Elm architecture. They must define:

fn init() -> State { ... }
fn update(state: State, msg: Msg) -> UpdateResult { ... }
fn view(state: State) -> UINode { ... }

Where State, Msg, Effect, UpdateResult, UINode, and PolicySet are the framework protocol types. See FRAMEWORK_SPEC.md for the full protocol.

Syntax quick reference

// Comments use double-slash

// Variables (type required; add `mut` to rebind later)
let x: Int = 42
let mut n: Int = 0
n = n + 1

// Loops
for item in [1, 2, 3] { n = n + item }
while n > 0 { n = n - 1 }

// Function
fn square(n: Int) -> Int {
    n * n
}

// Capability function
fn now() -> Int !{time.now} {
    // implementation
}

// Record literal
Point { x: 1, y: 2 }

// Record spread
Point { ..point, x: 10 }

// Enum variant
Shape::Circle(5.0)

// Pattern match
match x {
    0 => "zero"
    _ => "nonzero"
}

// Option
Some(42)
None

// Result
Ok("value")
Err("message")

// List
[1, 2, 3]

// Map
{ "key": "value" }

Import statements

import "std-name"

Import statements load a standard library package at compile time. The library source is inlined into the compilation unit before type-checking. The import line itself is removed from the compiled output.

Standard library packages are resolved from the libs/ directory relative to the current working directory. When a library source is inlined, any fn main() -> Int stub present in the library file is stripped so it does not conflict with the importing program’s own main.

Example:

import "std-json"

fn main() -> Int {
    let s: String = int_to_string(42)
    0
}

Built-in functions

These functions are provided by the runtime and do not need to be imported:

FunctionSignatureDescription
__builtin_int_to_string(Int) -> StringConvert an integer to its decimal string representation
__builtin_float_to_string(Float) -> StringConvert a float to its string representation
__builtin_string_len(String) -> IntLength of a string in bytes
__builtin_string_chars(String) -> List<String>Split a string into a list of single-character strings
__builtin_string_contains(String, String) -> BoolTrue if first string contains the second
__builtin_string_starts_with(String, String) -> BoolTrue if string starts with prefix
__builtin_string_ends_with(String, String) -> BoolTrue if string ends with suffix
__builtin_string_to_upper(String) -> StringUppercase copy
__builtin_string_to_lower(String) -> StringLowercase copy
__builtin_string_trim(String) -> StringStrip leading/trailing whitespace
__builtin_string_join(List<String>, String) -> StringJoin list with separator
__builtin_string_split(String, String) -> List<String>Split string on a delimiter
__builtin_string_replace(String, String, String) -> StringReplace first occurrence of pattern
__builtin_string_slice(String, Int, Int) -> StringSubstring by byte offsets
__builtin_int_parse(String) -> Result<Int, String>Parse a decimal integer string
__builtin_float_parse(String) -> Result<Float, String>Parse a float string
__builtin_bool_to_string(Bool) -> StringConvert a bool to “true” or “false”
__builtin_list_len(List<T>) -> IntNumber of elements
__builtin_list_is_empty(List<T>) -> BoolTrue if list has zero elements
__builtin_list_head(List<T>) -> Option<T>First element, or None
__builtin_list_tail(List<T>) -> List<T>All elements after the first
__builtin_list_append(List<T>, T) -> List<T>New list with item added at end
__builtin_list_concat(List<T>, List<T>) -> List<T>Concatenate two lists
__builtin_list_reverse(List<T>) -> List<T>Reversed copy
__builtin_map_get(Map<String, V>, String) -> Option<V>Look up a key; returns Some(v) or None
__builtin_map_set(Map<String, V>, String, V) -> Map<String, V>Return a new map with key set to value
__builtin_map_remove(Map<String, V>, String) -> Map<String, V>Return a new map with key removed
__builtin_map_contains_key(Map<String, V>, String) -> BoolTrue if key is present
__builtin_map_keys(Map<String, V>) -> List<String>All keys in sorted order
__builtin_map_values(Map<String, V>) -> List<V>All values in key-sorted order
__builtin_map_len(Map<String, V>) -> IntNumber of entries

These built-ins are also wrapped in std-json (via int_to_string, json_escape) and can be called directly in any .ax file.

Note on naming: The __builtin_ prefix distinguishes these from user-defined functions and prevents shadowing. User-facing wrappers in stdlib packages use cleaner names.

What .ax is not

.ax is deliberately minimal. It does not have:

  • Mutable variables (use record spread for state transitions)
  • Loops (use recursion or standard library functions)
  • Exceptions (use Result<T, E>)
  • Implicit side effects (every effect must be declared)
  • Generics (types are concrete at definition time)

These omissions are intentional. They keep the language deterministic and auditable.

MCP Server Tool Reference

The boruna-mcp binary exposes Boruna’s toolchain to AI coding agents (Claude Code, Cursor, Codex, …) over a JSON-RPC stdio transport. This page documents the wire contract for every tool — parameter names, types, return shapes, and error_kind values.

If you only want to register the server in your IDE, see AGENTS.md. The Rust source of record is crates/boruna-mcp/src/server.rs.

Quick start

# Register the server in your IDE's MCP config (e.g. .mcp.json for Claude Code):
{
  "mcpServers": {
    "boruna": {
      "command": "boruna-mcp",
      "args": ["--templates-dir", "/path/to/templates", "--libs-dir", "/path/to/libs"],
      "env": {}
    }
  }
}

Both --templates-dir and --libs-dir are optional. Defaults are templates and libs relative to the working directory.

Conventions

The tools below share several conventions:

  • All tools return JSON inside an MCP Content::text payload. Responses are pretty-printed.
  • Every response — success and failure — carries protocol_version: 1. This is the wire-format version of the response envelope. It bumps only on a breaking shape change (field rename, removal, type change, or error_kind semantics change). Additive changes keep the version. Integrators should reject any response whose protocol_version exceeds the version they were built against, and may safely upgrade their parsers when the field stays the same. See the Stability section for the full versioning policy.
  • Domain errors are returned as success: false JSON, not as MCP errors. This includes compile failures, runtime errors, validation errors, parse errors, etc. MCP-protocol errors (returned as McpError) are reserved for transport-level problems — most commonly when a source argument exceeds the 1 MB limit enforced by every source-accepting tool.
  • Source code is passed as a string, not a file path. Tool callers are responsible for reading files themselves.
  • Synchronous Boruna APIs run inside tokio::task::spawn_blocking. That means tool calls don’t starve the MCP event loop, but each call is single-threaded.
  • success: true responses always include the success flag. Field shapes after the flag vary per tool — see each section.
  • error_kind values are stable strings. Integrators may switch on them safely. New error_kind values may be added in a non-breaking way; existing ones are not renamed.

Tools

Note on the JSON examples below: for brevity, the per-tool examples show only the body fields. Every actual response — success and failure — also includes "protocol_version": 1 immediately after "success". See the Conventions and Stability sections for the full contract.

boruna_compile

Compile .ax source code and return module info.

Parameters

FieldTypeRequiredDescription
sourcestringyesThe .ax source code to compile. Max 1 MB.
namestringnoModule name. Default "module".

Returns (success)

{
  "success": true,
  "module": {
    "name": "module",
    "version": 1,
    "functions": 3,
    "types": 1,
    "constants": 7,
    "entry": 0
  }
}

Returns (failure)

{
  "success": false,
  "errors": [
    { "severity": "error", "code": "E001", "message": "...", "line": 10, "col": 5 }
  ]
}

Error codes: E001 (lexer), E002 (parser), E008 (codegen), E009 (typechecker).


boruna_ast

Parse .ax source code and return the AST as JSON.

Parameters

FieldTypeRequiredDescription
sourcestringyesThe .ax source code to parse. Max 1 MB.

Returns (success)

{
  "success": true,
  "truncated": false,
  "ast": { /* program AST as JSON */ }
}

If the AST exceeds 100 KB, the tool returns the truncated string and metadata instead of the parsed object:

{
  "success": true,
  "truncated": true,
  "ast_size": 134217,
  "ast": "..."
}

Returns (failure) — same errors shape as boruna_compile (lexer or parser errors only), or the following if AST JSON serialization itself fails (rare):

{ "success": false, "error_kind": "serialization_error", "message": "..." }

boruna_run

Compile and execute .ax source code under a capability policy.

Parameters

FieldTypeRequiredDescription
sourcestringyesThe .ax source code to run. Max 1 MB.
policystring | objectno"allow-all" / "deny-all" shorthand, or a Policy object — see policy-schema.md. Default "allow-all".
max_stepsinteger (u64)noVM step ceiling. Default 10000000.
tracebooleannoEmit an opcode-level execution trace in the response. Default false.

There is no input parameter. Script authors interpolate runtime values as literals into the .ax source before submission. This keeps the determinism contract clean: the (source, policy) tuple fully determines the run.

Returns (success)

{
  "success": true,
  "result": <value>,
  "steps": 142,
  "ui_output": [<value>, ...]
}

If trace: true:

{
  "success": true,
  "result": <value>,
  "steps": 142,
  "ui_output": [],
  "trace": ["op:0", "op:1", ...],
  "trace_truncated": false
}

The trace is capped at 500 entries; trace_truncated: true indicates the suffix was discarded.

Returns (failure)

{ "success": false, "error_kind": "runtime_error",  "message": "...", "steps": 7 }
{ "success": false, "error_kind": "invalid_policy", "message": "..." }

Plus the errors shape from boruna_compile if compilation fails before the run starts.

error_kind values: runtime_error, invalid_policy. Compile failures are returned without an error_kind (they use the errors array). Exceeding max_steps is reported as runtime_error with a step-limit message — there is no distinct kind for it.

Value encoding — primitives are passed through directly (Int → number, String → string, Bool → bool, Unit → null). Tagged values use the shapes:

{"option": "None"}                         // Value::None
{"option": "Some", "value": <inner>}       // Value::Some
{"result": "Ok",   "value": <inner>}       // Value::Ok
{"result": "Err",  "value": <inner>}       // Value::Err
{"type": "record", "type_id": 4, "fields": [...]}
{"type": "enum",   "type_id": 6, "variant": "Tag", "payload": <inner>}
{"actor_id": 1}                            // Value::ActorId
{"fn_ref": 12}                             // Value::FnRef

Value::List becomes a JSON array; Value::Map becomes a JSON object.

Progress notifications (formalized as 1.x stable in post1-T-1.1; underlying mechanism shipped in sprint 0.4-S6).

When the caller includes a progressToken in the request’s _meta field — the standard MCP mechanism for streaming progress — the server drives the VM through Vm::execute_bounded in slices of ~100,000 opcodes. Between slices it emits MCP notifications/progress events whose progress field is the cumulative VM step count:

{
  "jsonrpc": "2.0",
  "method": "notifications/progress",
  "params": {
    "progressToken": "<echoed from request>",
    "progress": 142000
  }
}

The events are coarse — slice-bounded, not per-opcode — so a long-running script emits several events per second on typical hardware. Clients that don’t supply a progressToken see no notifications and no extra latency. The streaming and non-streaming paths share their semantics for start_time, max_wall_ms, error handling, and final response shape, so progress-aware clients see identical results to legacy clients on the same input.


boruna_check

Run diagnostics on .ax source.

Parameters

FieldTypeRequiredDescription
sourcestringyesThe .ax source code. Max 1 MB.
file_namestringnoFilename used in diagnostic locations. Default "<source>".

Returns

{
  "success": true,
  "file": "<source>",
  "diagnostics_count": 2,
  "diagnostics": [
    {
      "id": "MISSING_MAIN",
      "severity": "error",
      "message": "...",
      "location": { "file": "<source>", "line": 1, "col": 1, "end_line": 1, "end_col": 1 },
      "patches": [
        {
          "id": "add_main",
          "description": "...",
          "confidence": "High",
          "rationale": "..."
        }
      ]
    }
  ]
}

location and patches are present only when applicable. confidence is one of "High", "Medium", "Low".

This tool returns success: true even when diagnostics are found — diagnostics are findings, not failures. Check diagnostics_count and per-entry severity to react.


boruna_repair

Auto-repair .ax source using diagnostic patches.

Parameters

FieldTypeRequiredDescription
sourcestringyesThe .ax source code to repair. Max 1 MB.
file_namestringnoFilename used in diagnostic locations. Default "<source>".
strategystringno"best" (default) — apply the highest-confidence patch per diagnostic. "all" — apply every patch. Ignored if patch_id is set.
patch_idstringnoApply only the patch with this ID; sets strategy to "by_id".

Returns

{
  "success": true,
  "repaired_source": "<the modified .ax source>",
  "patches_applied": 2,
  "patches_skipped": 0,
  "applied": [{ "diagnostic_id": "...", "patch_id": "...", "description": "..." }],
  "skipped": [{ "diagnostic_id": "...", "reason": "..." }],
  "verify_passed": true,
  "diagnostics_before": 2,
  "diagnostics_after": 0
}

verify_passed: true means re-running diagnostics on the repaired source produced no errors. Always inspect diagnostics_after before trusting the repair.


boruna_validate_app

Validate that .ax source conforms to the App protocol (Elm-style init / update / view).

Parameters

FieldTypeRequiredDescription
sourcestringyesThe .ax source code. Max 1 MB.

Returns

{
  "success": true,
  "has_init": true,
  "has_update": true,
  "has_view": true,
  "has_policies": false,
  "state_type": "State",
  "message_type": "Msg",
  "errors": [],
  "valid": true
}

valid: true when errors is empty AND all three of has_init / has_update / has_view are true.

Returns (failure) — success: false, error_kind: "framework_error" if the validator itself crashed; compile errors use the errors shape from boruna_compile.


boruna_framework_test

Run a framework App by sending a sequence of messages.

Parameters

FieldTypeRequiredDescription
sourcestringyesThe .ax framework app source. Max 1 MB.
messagesstring[]yesMessages as "tag:payload" strings (e.g. ["increment:1", "reset:0"]). Payloads parse as integer if possible, otherwise string.

Returns

{
  "success": true,
  "init_state": <value>,
  "cycles": [
    { "message": "increment:1", "state": <value>, "effects": 0, "ui_tree": <value or null> },
    ...
  ],
  "final_state": <value>,
  "total_cycles": 2
}

If a cycle fails, the tool returns success: false, error_kind: "framework_error", and includes the partial cycles (with the failing entry containing an error field) plus the init_state so callers can debug the divergence.

Value formatting in this tool is brief (different from boruna_run — see source format_value_brief): records render as flat field arrays, enums as {variant, payload}, options/results as {Tag: value}.


boruna_workflow_validate

Validate a workflow definition (JSON).

Parameters

FieldTypeRequiredDescription
workflow_jsonstringyesThe full workflow.json content as a string.

Returns

{
  "success": true,
  "workflow_name": "code_review",
  "workflow_version": "1.0.0",
  "steps_count": 3,
  "edges_count": 2,
  "execution_order": ["fetch_pr", "analyze", "post_comment"]
}

Returns (failure)

{ "success": false, "error_kind": "parse_error",      "message": "..." }
{ "success": false, "error_kind": "validation_error", "errors": [{ "kind": "Cycle", "message": "..." }] }

error_kind: "parse_error" means the JSON itself was malformed. validation_error means the JSON parsed but the DAG is invalid (cycle, missing step reference, etc.). Each entry in errors carries a kind (debug-formatted Rust enum variant) and a human-readable message.


boruna_template_list

List available Boruna app templates.

Parameters — none.

Returns

{
  "success": true,
  "count": 5,
  "templates": [
    {
      "name": "crud-admin",
      "version": "1.0.0",
      "description": "...",
      "dependencies": ["std-ui", "std-forms"],
      "capabilities": ["db.query"],
      "args": ["entity_name", "fields"]
    }
  ]
}

args lists only the variable names; consult template apply for substitution.

Returns (failure) — success: false, error_kind: "template_error", message: ... (e.g. when the templates dir doesn’t exist or contains malformed manifests).


boruna_template_apply

Apply a template with variable substitution.

Parameters

FieldTypeRequiredDescription
template_namestringyesTemplate name (e.g. "crud-admin").
argsstring[]yesArguments as "key=value" strings (e.g. `[“entity_name=products”, “fields=name
validatebooleannoCompile the rendered output to verify it parses. Default false.

Returns

{
  "success": true,
  "template_name": "crud-admin",
  "output_file": "...",
  "source": "<rendered .ax source>",
  "dependencies": ["std-ui"],
  "capabilities": ["db.query"],
  "validation": { "passed": true }
}

The validation object is present only when validate: true. If validation fails:

"validation": { "passed": false, "error": "..." }

Returns (failure)

{ "success": false, "error_kind": "invalid_args",   "message": "argument must be key=value format, got: ..." }
{ "success": false, "error_kind": "template_error", "message": "..." }

Limits

  • Source size: every tool that accepts a source parameter rejects payloads above 1 MB at the MCP layer (returned as an MCP invalid_params error, not as JSON). This is enforced in crates/boruna-mcp/src/server.rs::validate_source.
  • AST size (boruna_ast): ASTs above 100 KB are returned truncated as a string (with truncated: true and ast_size), not as a parsed JSON object.
  • Trace size (boruna_run): execution traces are capped at 500 entries (trace_truncated indicates suffix discarded).
  • Process model: all tool calls run synchronously inside spawn_blocking. The MCP server is single-tenant by design; long-running tool calls block the response, not the event loop.

Stability

  • protocol_version: 1 is the wire-format version of the response envelope, present on every tool response (success and failure). Locked by crates/boruna-mcp/src/tools/mod.rs::TOOL_RESPONSE_PROTOCOL_VERSION; a regression test (protocol_version_tests) asserts coverage across both success and failure paths of every tool. Bumped only on a breaking shape change anywhere in the envelope; additive changes keep the version.
  • Tool names (boruna_compile, boruna_run, …) are stable. Renames require a major version bump.
  • error_kind values are stable strings. New ones may be added; existing ones are not renamed.
  • Top-level response fields (success, protocol_version, named result fields) are stable. New fields are additive.
  • Value encoding inside result (the tagged shapes for Option / Result / records / enums) is locked.

Versioning policy for protocol_version

ChangeBump?
Adding a new optional response fieldNo
Adding a new error_kind valueNo
Adding a new toolNo
Renaming an existing fieldYes
Removing an existing fieldYes
Changing the type of an existing fieldYes
Changing what an existing error_kind meansYes
Changing the result value encoding (Option/Result/record tagged shapes)Yes

When protocol_version bumps, integrators see protocol_version: 2 (or higher) and can branch on the version to support both shapes during a migration window.

Pairs with Policy.schema_version (currently 1) — together they cover both the request-side policy schema and the response-side envelope.

See also

Capability Policy Schema

The policy parameter on the MCP boruna_run tool (and the --policy <file> flag on the boruna CLI) accepts either:

  • A string shorthand: "allow-all" or "deny-all"
  • A Policy object matching the schema below

This page documents the object form. The machine-readable schema lives at policy.schema.json.

Object form

{
  // Policy schema version. Currently always 1. Optional.
  "schema_version": 1,

  // Default behavior for capabilities NOT listed in `rules`.
  // false = deny by default (allowlist mode); true = allow by default (denylist mode).
  // Required for predictable behavior — do not omit.
  "default_allow": false,

  // Per-capability rules. Keys are capability names (see table below).
  "rules": {
    "net.fetch": {
      "allow":  true,   // boolean, required
      "budget": 0       // u64, required. 0 = unlimited; otherwise hard ceiling on call count.
    }
  },

  // Optional network-specific controls. Applied when `net.fetch` is allowed.
  "net_policy": {
    "allowed_domains":      ["api.openai.com", "*.our-api.example"], // empty = all
    "allowed_methods":      ["GET", "POST"],                          // empty = all
    "max_response_bytes":   10485760,                                 // default 10 MB
    "timeout_ms":           30000,                                    // default 30 s
    "allow_redirects":      true                                      // default true
  }
}

Capability names

These are the strings you use as keys in rules. They mirror boruna_bytecode::Capability::name().

CapabilityKeyNotes
Network fetchnet.fetchHTTP GET/POST/etc. — also gated by net_policy
Filesystem readfs.read
Filesystem writefs.write
Database querydb.query
UI renderui.renderFramework view() output
Current timetime.nowNon-deterministic; deny in pure pipelines
Random numberrandomNon-deterministic; deny in pure pipelines
LLM callllm.callExternal model invocation — apply budget to cap cost
Spawn actoractor.spawn
Send to actoractor.send

The strict validator rejects aliases. Sprint 0.4-S15 locked the rule-key surface to canonical names only. A policy file with "net" as a rule key fails validation with error_kind: "policy.invalid_capability" and a hint to use "net.fetch". Aliases were silently no-ops at gateway-check time before — fixing that footgun was the point of 0.4-S15 (project convention #1: reject at parse, don’t silently override).

Examples

1. Allowlist domain only — deny everything except net.fetch to api.openai.com

{
  "default_allow": false,
  "rules": { "net.fetch": { "allow": true, "budget": 0 } },
  "net_policy": { "allowed_domains": ["api.openai.com"] }
}

2. Allow-all minus filesystem writes — useful for read-only workflows

{
  "default_allow": true,
  "rules": { "fs.write": { "allow": false, "budget": 0 } }
}

3. LLM call quota — cap LLM invocations at 5 per run

{
  "default_allow": true,
  "rules": { "llm.call": { "allow": true, "budget": 5 } }
}

When the budget is exceeded the run aborts with a runtime_error whose message references CapabilityBudgetExceeded(LlmCall).

Surprising behavior to know

  • default_allow defaults to false. A Policy {} (empty object) denies everything. Always set default_allow explicitly.
  • budget: 0 means unlimited, not “zero allowed.” Use { "allow": false, "budget": 0 } to deny.
  • String shorthand and object form are not mixable. Pass exactly one shape.
  • Unknown JSON shapes are rejected. Old MCP clients that accidentally posted typo’d strings (e.g. "alow-all") used to be silently treated as allow-all. They now return success: false, error_kind: "invalid_policy". This is intentional — silent fall-through to allow-all was the bug FleetQ reported.
  • Unknown fields are rejected (sprint 0.4-S15). A typo like "default_alow": true no longer parses as default_allow: false (silent default); it fails with error_kind: "policy.unknown_field".
  • Unsupported schema_version values are rejected. This binary supports schema_version: 1. Setting 2 or any other value fails with error_kind: "policy.unknown_schema_version".

CLI tooling (sprint 0.4-S15)

# Strict-validate a policy file. Designed as a CI gate.
boruna policy validate policies/prod.json
# → exit 0 + "OK: ..." on success
# → exit 2 + stderr "error: policy.<kind>: ..." on validation error
# → exit 1 + stderr "error: policy.io_error: ..." on file IO error

# Machine-parseable output:
boruna policy validate --json policies/prod.json
# → {"ok":true} or {"ok":false,"errors":[{"error_kind":"policy.unknown_field",...}]}

# Print the effective policy (denormalized).
boruna policy show policies/prod.json

The MCP server exposes the same validator as boruna_policy_validate.

Stable error_kind taxonomy

The strict validator emits these stable strings (project convention #2 — locked forever):

error_kindWhen
policy.io_errorFile missing or unreadable
policy.parse_errorJSON syntax error or value type mismatch
policy.unknown_schema_versionschema_version is set to an unsupported value
policy.unknown_fieldUnknown field at any level (top-level, net_policy, or inside a rule)
policy.invalid_capabilityA rule key is not a recognized canonical capability name
policy.invalid_net_policynet_policy value out of range or unknown method

The boruna_run MCP tool also emits the legacy error_kind: "invalid_policy" for non-object input (string typos, arrays, numbers). The new policy.* kinds apply to object-form payloads only — they are additive over invalid_policy, not a replacement.

Versioning

The schema carries schema_version: 1. Future breaking changes will bump this number; the MCP tool will continue to accept the old shape as long as schema_version matches a supported value. This field is what lets you cache (script_hash, policy_hash) results safely across binary upgrades. Sprint 0.4-S15 locked this contract: only 1 is currently accepted; new optional fields can be added at v1; shape changes require a version bump.

Hashing for caching

Because Policy is Serialize + Deserialize, you can hash a normalized policy for cache keys:

#![allow(unused)]
fn main() {
let bytes = serde_json::to_vec(&policy).unwrap();
let hash  = sha2::Sha256::digest(&bytes);
}

Pair hash(policy) with hash(source) to memoize deterministic runs. (The capability-set identity portion — making the hash stable across binary upgrades — is tracked separately in the project roadmap.)

Capability Identity & Caching Contract

Boruna binaries advertise a stable, hashable identity for their capability surface. Integrators use this identity to safely cache deterministic run results across binary upgrades.

Stability: stable from 0.3.0. Implementation shipped in 0.2.x for early integrators (FleetQ #3).

Why this exists

Boruna is deterministic by construction: for a given (source, inputs, policy, capability_contract) the run result is identical every time. That means an integrator can memoize results indefinitely — as long as the capability contract hasn’t changed.

Without a stable identity for the capability contract, integrators have two bad options:

  1. Don’t cache — leave free determinism on the table.
  2. Cache by binary version string — invalidate on every patch release, even when contracts didn’t change.

capability_set_hash solves this: it’s a content-addressed identity over the capability surface that changes only when the contract changes.

Surface

CLI

$ boruna capability list --json
{
  "protocol_version": 1,
  "name": "boruna",
  "version": "0.2.0",
  "capabilities": [
    { "name": "actor.send",   "version": "1" },
    { "name": "actor.spawn",  "version": "1" },
    { "name": "db.query",     "version": "1" },
    { "name": "fs.read",      "version": "1" },
    { "name": "fs.write",     "version": "1" },
    { "name": "llm.call",     "version": "1" },
    { "name": "net.fetch",    "version": "1" },
    { "name": "random",       "version": "1" },
    { "name": "time.now",     "version": "1" },
    { "name": "ui.render",    "version": "1" }
  ],
  "capability_set_hash": "sha256:b0ca1793a79656447d560092bae7af4b0ebee82023c6d2bea82bd80621bd9637"
}

The CLI conveys success via process exit code, so there is no success field.

MCP

Tool boruna_capability_list (no parameters). Returns the same fields as the CLI --json output plus a leading success: true envelope (this server’s universal convention for tool responses):

{
  "success": true,
  "protocol_version": 1,
  "name": "boruna",
  "version": "0.2.0",
  "capabilities": [ ... ],
  "capability_set_hash": "sha256:..."
}

Field semantics

FieldMeaningStability
protocol_versionWire-format version of this report. Bumped on breaking shape changes (field rename, removal, type change). Additive changes keep the version.Frozen for protocol_version: 1 going forward.
nameBinary identity. Defaults to "boruna" for upstream binaries. Downstream forks that rebrand may emit their own name. Does NOT participate in capability_set_hash.Stable string.
versionBinary version (Cargo.toml package version of the calling crate). Does NOT participate in capability_set_hash.Semver.
capabilities[].nameCapability identifier (e.g. "net.fetch").Stable; new caps appear, never rename.
capabilities[].versionCapability contract version. Bumped on argument/return/semantics changes.Increments as integer string.
capability_set_hashSHA-256 over canonical encoding of (name, version) pairs in sorted order.Algorithm frozen — see below.

Hash algorithm

The capability_set_hash is computed byte-for-byte as follows:

  1. Take all capabilities in canonical order (sorted ascending by name).
  2. For each, encode the UTF-8 bytes of "{name}\t{version}\n".
    • \t = ASCII 0x09 (tab)
    • \n = ASCII 0x0A (newline)
  3. Concatenate all encodings into a single byte string (no separators between capabilities beyond the trailing \n of each).
  4. SHA-256 of that byte string.
  5. Lower-case hex, prefixed with "sha256:".

Worked example (current 0.2.0 surface)

The byte string fed to SHA-256 is exactly:

actor.send\t1\nactor.spawn\t1\ndb.query\t1\nfs.read\t1\nfs.write\t1\nllm.call\t1\nnet.fetch\t1\nrandom\t1\ntime.now\t1\nui.render\t1\n

(252 bytes, with literal tabs and newlines, not the escape sequences shown.)

Reproduce with shell:

$ printf 'actor.send\t1\nactor.spawn\t1\ndb.query\t1\nfs.read\t1\nfs.write\t1\nllm.call\t1\nnet.fetch\t1\nrandom\t1\ntime.now\t1\nui.render\t1\n' | shasum -a 256
b0ca1793a79656447d560092bae7af4b0ebee82023c6d2bea82bd80621bd9637  -

Caching contract for integrators

Recommended cache key:

key = sha256(
  source_hash       ||  // sha256(.ax source bytes)
  policy_hash       ||  // sha256(canonical JSON of the Policy object)
  capability_set_hash || // from this endpoint
  policy_schema_version  // from the Policy.schema_version field, e.g. "1"
)

This guarantees:

  • A .ax source change invalidates the entry.
  • A policy change invalidates the entry.
  • A capability contract change invalidates the entry (new capability added, or existing capability semantics changed).
  • A policy schema change invalidates the entry (Boruna evolves the policy format).

Anything outside this set — Rust toolchain version, Boruna patch version that didn’t touch capabilities, build host — does not invalidate cached results, because none of it can change the deterministic output.

Per-capability version semantics

Each capability has its own version. We bump it only when the contract changes in a way that would make a downstream cached result invalid:

ChangeBump?
Argument shape changes (new field, removed field, type change)Yes
Return shape changesYes
Side-effect semantics change (e.g. net.fetch starts following redirects by default)Yes
Performance improvement, internal refactorNo
Bug fix that brings behavior in line with documented contractNo
Stricter input validation that rejects previously-accepted-but-undefined inputsJudgment call — usually Yes to be safe

When you bump a capability version, you must also:

  1. Update the match arm in crates/llmbc/src/capability.rs::Capability::version().
  2. Update the golden hash in crates/llmbc/src/tests.rs::test_capability_set_hash_known_value.
  3. Add a ### Changed entry under [Unreleased] in CHANGELOG.md referencing the capability.

What does and does not affect the hash

Affects capability_set_hash?
✅ Adding a new capabilityYes — extends the byte-string input.
✅ Removing a capabilityYes — removes from input.
✅ Bumping any capability’s version fieldYes — that’s the entire point.
❌ Bumping the binary’s version fieldNo — binary_version is metadata, not contract surface.
❌ Forks emitting a different nameNo — binary_name is metadata, not contract surface.
❌ Bumping protocol_versionNo — wire-format envelope is independent of capability contract.

This separation is what makes the hash useful: a Boruna patch release that touches no capabilities produces an identical hash, so cached results stay valid.

Stability guarantees

  • The algorithm above is frozen. We will never change how the hash is computed without a major version bump (protocol_version: 2) and a clearly-documented migration path. An algorithm-level test (test_compute_capability_set_hash_algorithm_known_value) locks the encoding rule independently of the current capability set.
  • The JSON shape of the report is locked at protocol_version: 1. Field additions are non-breaking and keep protocol_version; field removals, renames, or type changes bump protocol_version.
  • The per-capability versions evolve independently of each other and of protocol_version. Today they are all "1". Future bumps follow the rules above.

Pairs with Policy.schema_version

The Policy object (see policy schema) carries a schema_version field independently. Together they cover:

  • capability_set_hash — does the binary still mean the same thing by net.fetch?
  • Policy.schema_version — does the binary still parse my policy the same way?

Both must match for a cached result to be valid.

See also

Diagnostic Codes

Every diagnostic Boruna’s toolchain emits carries a stable E0NN code. Codes are stable forever — never reused, never renumbered. Tools and AI agents may switch on them.

The registry is machine-readable. Query it directly:

boruna lang codes          # human table
boruna lang codes --json   # { "version": 1, "codes": [ ... ] }

Codes appear in boruna lang check --json output as the id field of each diagnostic.

CodeNameCategorySummary
E001lexer-errorlexicalThe source could not be tokenized (invalid character or token).
E002parse-errorsyntaxThe token stream did not form a valid syntax tree.
E003undefined-variablename-resolutionA referenced variable is not defined in scope.
E004undefined-functionname-resolutionA called function is not defined in the module.
E005non-exhaustive-matchpattern-matchingA match expression does not cover all possible cases.
E006unknown-fieldtypeA record field access or construction references an unknown field.
E007capability-violationcapabilityA function performs an effect it does not declare in its capability set.
E008codegen-errorcodegenThe typechecked program could not be lowered to bytecode.
E009type-errortypeAn expression’s type does not match the type required by its context.
E010assign-to-immutabletypeA binding declared without mut (or a parameter or loop variable) is reassigned. A warning today; an error in language version 2.0.

The table above is generated from the same registry the CLI serves (tooling/src/diagnostics/registry.rs). A drift test asserts the registry stays 1:1 with the E0NN constants the compiler emits.

Structured Diagnostics + Suggested Patches

Overview

The tooling crate provides machine-readable diagnostics and auto-repair for .ax source files.

Every diagnostic is emitted in two formats:

  • Human-readable text (one line per diagnostic)
  • Stable JSON (version 1, serializable as diagnostics.json)

Error Codes

CodeCategoryDescription
E001LexerInvalid token / unexpected character
E002ParseSyntax error
E003TypeUndefined variable (with name suggestion)
E004TypeUndefined function
E005AnalysisNon-exhaustive match (missing enum variants)
E006AnalysisUnknown record field (with closest-name suggestion)
E007AnalysisCapability violation (update/view must be pure)
E008CodegenCode generation error
E009TypeGeneral type error
E010AnalysisBinding declared without mut (or a parameter / loop variable) is reassigned. Warning today, error in language version 2.0; lang repair adds mut

Suggested Patches

Each diagnostic may include suggested_patches. Each patch contains:

  • id: Stable identifier for this fix
  • description: Human-readable summary
  • confidence: high, medium, or low
  • rationale: Why this fix is suggested
  • edits: Array of TextEdit (file, start_line, old_text, new_text)

TextEdits are compatible with the PatchBundle Hunk format.

JSON Format

{
  "version": 1,
  "file": "path/to/file.ax",
  "diagnostics": [
    {
      "id": "E005",
      "severity": "error",
      "message": "non-exhaustive match on 'msg' of type 'Msg': missing variants: Clear, Remove",
      "location": {
        "file": "path/to/file.ax",
        "line": 7,
        "col": null,
        "end_line": null,
        "end_col": null
      },
      "suggested_patches": [
        {
          "id": "E005-add-arms",
          "description": "add missing match arms: Clear, Remove",
          "confidence": "high",
          "rationale": "match expression does not cover variants: Clear, Remove",
          "edits": [
            {
              "file": "path/to/file.ax",
              "start_line": 10,
              "old_text": "    }",
              "new_text": "        Clear => { /* TODO */ }\n        Remove => { /* TODO */ }\n    }"
            }
          ]
        }
      ],
      "related": []
    }
  ]
}

CLI Usage

Check

boruna lang check path/to/file.ax            # human-readable output
boruna lang check path/to/file.ax --json     # JSON to stdout
boruna lang check path/to/file.ax -o diag.json  # JSON to file

Repair

boruna lang repair path/to/file.ax                     # auto-fix with best suggestions
boruna lang repair path/to/file.ax --apply all          # apply all suggestions
boruna lang repair path/to/file.ax --apply E005-add-arms  # apply specific fix
boruna lang repair path/to/file.ax --from diag.json     # use pre-computed diagnostics

Analysis Passes

Match Exhaustiveness (E005)

Detects when a match on a typed enum parameter does not cover all variants. Skips if a wildcard (_) or catch-all identifier pattern is present.

Record Field Validation (E006)

Detects when a record literal uses a field name not present in the type definition. Suggests the closest valid field name via Levenshtein distance.

Capability Purity (E007)

In framework apps (those with init/update/view), detects when update() or view() declare capabilities. Suggests removing the !{...} annotation.

Undefined Variable (E003)

Enhances compiler errors with name suggestions. Collects all defined names (functions, parameters, local variables, types, builtins) and finds the closest match.

Repair Tool

The repair tool:

  1. Reads diagnostics (from JSON or runs check)
  2. Selects patches based on strategy (best/all/specific ID)
  3. Applies text edits to the source (reverse line order to avoid offset drift)
  4. Re-runs diagnostics to verify the fix
  5. Reports before/after diagnostic count and verify status

Patches are applied deterministically: same input always produces the same output.

Canonical error_kind taxonomy

This is the single source of truth for the stable error_kind strings emitted by the Boruna binary and the MCP server.

Stability contract

  • These strings are stable per docs/lts.md §B.6 (“Error taxonomy”). Once shipped in a tag, an error_kind is never renamed or removed inside the 1.x line.
  • New error_kind values MAY be added in 1.x minor releases; integrators MUST tolerate values they don’t recognize.
  • Integrators MAY switch on these strings programmatically — the strings are part of the LTS-protected contract, not human-readable log copy.

How this list is maintained

Every entry below is verified against a literal-string grep of the crates/ and orchestrator/ trees. When a new error_kind is added to the source, it MUST be added here in the same change. CI may grow a gate for this in a future sprint; today the discipline is reviewer- enforced.


evidence.* — evidence bundle reader

Emitted by boruna evidence verify and boruna evidence inspect. See orchestrator/src/audit/encryption.rs::EncryptionError for the source strings.

error_kindPhaseWhere it firesSprintCaller-facing meaning
evidence.encryption_key_requiredN/Aorchestrator/src/audit/verify.rs::verify_bundleW6-BBundle is encrypted (manifest carries an encryption block) but no KEK has been supplied via --bundle-encryption-key or BORUNA_BUNDLE_KEK.
evidence.encryption_key_mismatchN/Aorchestrator/src/audit/verify.rs::verify_bundleW6-BSupplied KEK does not unwrap the bundle’s wrapped_dek (wrong key, or the DEK ciphertext was tampered).
evidence.cipher_tag_invalidN/Aorchestrator/src/audit/verify.rs::verify_bundleW6-BAES-GCM authentication tag failed for at least one encrypted file — the bundle has been tampered after recording. Plaintext bytes are not returned to the caller.
evidence.unsupported_algorithmN/A(reserved)W6-Bencryption.algorithm is set to a value other than "aes-256-gcm". Reserved string for forward-compat per docs/spec/evidence-bundle-1.0.md §8.1.

Note on the W1-C reader gate. Bundles missing bundle.json or carrying an incompatible major format_version are rejected by verify_bundle with the diagnostic unsupported evidence bundle format_version: found '<x>', expected major '<y>'. This message is emitted as a VerifyError, not as a JSON error_kind field; tools wrapping the reader translate it to their own taxonomy. See orchestrator/src/audit/verify.rs::VerifyError.

workflow.* — workflow JSON definition reader

Emitted by boruna_orchestrator::WorkflowDef::from_json and surfaces in boruna workflow validate / boruna workflow run / the coord POST /api/runs path.

error_kindPhaseWhere it firesSprintCaller-facing meaning
workflow.missing_schema_versionserializationorchestrator/src/workflow/definition.rs::DefinitionError::error_kindW4workflow.json has no schema_version field. Required since v1.0; legacy workflows must be migrated.
workflow.unsupported_schema_versionserializationorchestrator/src/workflow/definition.rs::DefinitionError::error_kindW4workflow.json carries a schema_version value this binary doesn’t accept (e.g. 2 on a 1.x binary).
workflow.invalid_jsonserializationorchestrator/src/workflow/definition.rs::DefinitionError::error_kindW4workflow.json is not valid JSON or fails the workflow schema after the version gate.

policy.* — policy schema validator

Emitted by boruna policy validate and boruna_run (object-form policy input). See docs/reference/policy-schema.md for full context.

error_kindPhaseSprintCaller-facing meaning
policy.io_errorserialization0.4-S15Policy file missing or unreadable.
policy.parse_errorserialization0.4-S15JSON syntax error or value-type mismatch.
policy.unknown_schema_versionserialization0.4-S15schema_version set to an unsupported value.
policy.unknown_fieldserialization0.4-S15Unknown field at any level (top-level, net_policy, or inside a rule).
policy.invalid_capabilityserialization0.4-S15Rule key is not a recognized canonical capability name (aliases like net are rejected).
policy.invalid_net_policyserialization0.4-S15net_policy value out of range or unknown HTTP method.

MCP-layer top-level kinds

Emitted by the boruna-mcp server’s tool layer. These predate the namespaced evidence.* / workflow.* schemes and are kept for back-compat per the LTS contract.

error_kindToolPhaseSprintCaller-facing meaning
invalid_policyboruna_runserialization0.2.0Non-object policy input (string typo, array, number) was supplied. Object-form input that fails strict validation surfaces as a policy.* kind instead.
invalid_output_schemaboruna_runserialization0.4-S16The supplied output JSON-schema is malformed or the run’s output does not validate against it.
unsupported_limitboruna_runserialization0.4-S15A limits.* field is set to a value this binary cannot enforce yet.
parse_errorboruna_workflow_validate, boruna_compileserialization0.2.0Input JSON / source could not be parsed at the lexer or serde stage.
serialization_errorboruna_compileserialization0.2.0AST or compile output could not be serialized for return; internal-encoding failure.
validation_errorboruna_workflow_validateoutput_validation0.2.0Workflow JSON parsed but failed structural validation (cycle, missing field, unknown step reference).
validation_failedboruna_runoutput_validation0.4-S16Run output failed JSON-schema validation. Response body carries per-path errors.
runtime_errorboruna_runexecution0.2.0VM error during execution — capability denied, type mismatch, etc. The error field carries the message.
limit_exceededboruna_runexecution / serialization0.4-S15A configured limit was hit. limit_kind discriminates: step_limit, wall_ms (execution), output_bytes (serialization).
framework_errorboruna_validate_app, boruna_framework_testexecution0.2.0Framework App protocol validation or test-harness error (init/update/view shape mismatch, message dispatch failure).
template_errorboruna_template_applyexecution0.2.0Template substitution failed (missing variable, unknown template, manifest-validation failure at apply time).
invalid_argsboruna_template_applyserialization0.2.0Template --args payload could not be parsed as key=value pairs.

Conventions

  • All error_kind strings are dotted, lower-snake-case, and hierarchical (<namespace>.<short_kind>). The namespace identifies the surface (evidence reader, workflow loader, policy validator, MCP top-level).
  • “Phase” follows the project convention of distinguishing serialization (parse-time / shape rejection) from output_validation (post-execution shape rejection) from execution (runtime failures). N/A means the kind is a control-flow / policy-gate decision rather than a shape error.

Cross-references

Framework Public API

Stability: Experimental within the Boruna 1.x line. The framework is the Elm-architecture runtime exposed by boruna_framework; it ships with the rest of Boruna under the workspace version (currently v1.0.0-rc2) but its public types are NOT included in the LTS contract §B. Expect API changes between Boruna 1.x minor releases as adoption feedback comes in. See stability.md for the tier definitions and lts.md for the surfaces that ARE LTS-protected.

All unlisted types/functions are internal and may change without notice.

boruna_framework (crate root re-exports)

#![allow(unused)]
fn main() {
pub use error::FrameworkError;
pub use runtime::AppRuntime;
pub use validate::AppValidator;
pub use testing::TestHarness;
pub use policy::PolicySet;
}

boruna_framework::error

#![allow(unused)]
fn main() {
pub enum FrameworkError {
    Validation(String),
    MissingFunction(String),
    PurityViolation { name: String },
    WrongArity { name: String, expected: usize, got: usize },
    MissingType(String),
    Effect(String),
    PolicyViolation(String),
    State(String),
    Compile(boruna_compiler::CompileError),
    Runtime(boruna_vm::VmError),
    MaxCyclesExceeded(u64),
}
}

boruna_framework::validate

#![allow(unused)]
fn main() {
pub struct AppValidator;

impl AppValidator {
    pub fn validate(program: &Program) -> Result<ValidationResult, FrameworkError>;
    pub fn is_valid_app(program: &Program) -> bool;
}

pub struct ValidationResult {
    pub has_init: bool,
    pub has_update: bool,
    pub has_view: bool,
    pub has_policies: bool,
    pub state_type: Option<String>,
    pub message_type: Option<String>,
    pub errors: Vec<String>,
}
}

boruna_framework::runtime

#![allow(unused)]
fn main() {
pub struct AppMessage {
    pub tag: String,
    pub payload: Value,
}

impl AppMessage {
    pub fn new(tag: impl Into<String>, payload: Value) -> Self;
    pub fn to_value(&self) -> Value;
}

pub struct CycleRecord {
    pub cycle: u64,
    pub message: AppMessage,
    pub state_before: Value,
    pub state_after: Value,
    pub effects: Vec<Effect>,
    pub ui_tree: Option<Value>,
}

pub struct AppRuntime { /* private fields */ }

impl AppRuntime {
    pub fn new(module: Module) -> Result<Self, FrameworkError>;
    pub fn state(&self) -> &Value;
    pub fn cycle(&self) -> u64;
    pub fn cycle_log(&self) -> &[CycleRecord];
    pub fn policy(&self) -> &PolicySet;
    pub fn state_machine(&self) -> &StateMachine;
    pub fn send(&mut self, msg: AppMessage) -> Result<(Value, Vec<Effect>, Option<Value>), FrameworkError>;
    pub fn view(&self) -> Result<Value, FrameworkError>;
    pub fn snapshot(&self) -> String;
    pub fn rewind(&mut self, cycle: u64) -> Result<(), FrameworkError>;
    pub fn diff_from(&self, cycle: u64) -> Vec<StateDiff>;
}
}

boruna_framework::effect

#![allow(unused)]
fn main() {
pub struct Effect {
    pub kind: EffectKind,
    pub payload: Value,
    pub callback_tag: String,
}

pub enum EffectKind {
    HttpRequest, DbQuery, FsRead, FsWrite,
    Timer, Random, SpawnActor, EmitUi,
}

impl EffectKind {
    pub fn from_str(s: &str) -> Option<Self>;
    pub fn capability_name(&self) -> &'static str;
    pub fn as_str(&self) -> &'static str;
}

pub fn parse_effects(effects_value: &Value) -> Vec<Effect>;
pub fn parse_update_result(value: &Value) -> Option<(Value, Vec<Effect>)>;
}

boruna_framework::state

#![allow(unused)]
fn main() {
pub struct StateSnapshot {
    pub cycle: u64,
    pub state: Value,
    pub json: String,
}

pub struct StateDiff {
    pub field_index: usize,
    pub field_name: String,
    pub old_value: Value,
    pub new_value: Value,
}

pub struct StateMachine { /* private fields */ }

impl StateMachine {
    pub fn new(initial_state: Value) -> Self;
    pub fn current(&self) -> &Value;
    pub fn cycle(&self) -> u64;
    pub fn history(&self) -> &[StateSnapshot];
    pub fn transition(&mut self, new_state: Value);
    pub fn snapshot(&self) -> String;
    pub fn restore(&mut self, json: &str) -> Result<(), FrameworkError>;
    pub fn diff_from_cycle(&self, cycle: u64) -> Vec<StateDiff>;
    pub fn diff_values(old: &Value, new: &Value) -> Vec<StateDiff>;
    pub fn rewind(&mut self, target_cycle: u64) -> Result<(), FrameworkError>;
}
}

boruna_framework::ui

#![allow(unused)]
fn main() {
pub struct UINode {
    pub tag: String,
    pub props: Vec<(String, Value)>,
    pub children: Vec<UINode>,
}

impl UINode {
    pub fn new(tag: impl Into<String>) -> Self;
    pub fn with_prop(self, key: impl Into<String>, value: Value) -> Self;
    pub fn with_child(self, child: UINode) -> Self;
}

pub fn value_to_ui_tree(value: &Value) -> UINode;
pub fn ui_tree_to_value(node: &UINode) -> Value;
}

boruna_framework::policy

#![allow(unused)]
fn main() {
pub struct PolicySet {
    pub capabilities: Vec<String>,
    pub max_effects_per_cycle: u64,
    pub max_steps: u64,
}

impl PolicySet {
    pub fn allow_all() -> Self;
    pub fn from_value(value: &Value) -> Self;
    pub fn check_effect(&self, effect: &Effect) -> Result<(), FrameworkError>;
    pub fn check_batch(&self, effects: &[Effect]) -> Result<(), FrameworkError>;
}
}

boruna_framework::testing

#![allow(unused)]
fn main() {
pub struct TestHarness { /* private fields */ }

impl TestHarness {
    pub fn from_source(source: &str) -> Result<Self, FrameworkError>;
    pub fn state(&self) -> &Value;
    pub fn cycle(&self) -> u64;
    pub fn send(&mut self, msg: AppMessage) -> Result<(Value, Vec<Effect>), FrameworkError>;
    pub fn simulate(&mut self, messages: Vec<AppMessage>) -> Result<Value, FrameworkError>;
    pub fn assert_state_field(field_index: usize, expected: &Value) -> Result<(), FrameworkError>;
    pub fn assert_effects(expected_kinds: &[&str]) -> Result<(), FrameworkError>;
    pub fn assert_state(expected: &Value) -> Result<(), FrameworkError>;
    pub fn cycle_log(&self) -> &[CycleRecord];
    pub fn snapshot(&self) -> String;
    pub fn rewind(&mut self, cycle: u64) -> Result<(), FrameworkError>;
    pub fn replay_verify(&self, source: &str, messages: Vec<AppMessage>) -> Result<bool, FrameworkError>;
    pub fn view(&self) -> Result<Value, FrameworkError>;
    pub fn runtime(&self) -> &AppRuntime;
}

pub fn simulate_messages(source: &str, messages: Vec<AppMessage>) -> Result<Value, FrameworkError>;
}

Compliance Workflow Templates

Pre-built workflow patterns for common regulated use cases. Each template demonstrates how Boruna’s evidence bundle, capability policy, and approval gate features satisfy specific compliance requirements.

TemplateStandardKey Feature
soc2_audit_workflowSOC 2Hash-chained evidence bundle as tamper-evident audit trail
hipaa_data_pipelineHIPAAPHI redaction before evidence bundle write
financial_review_pipelineSOX / dual-controlMulti-approver approval gates with immutable sign-off record

How to use

  1. Copy the template to your project
  2. Replace synthetic data with real capability calls (e.g., net.fetch for live system data)
  3. Run with --record to produce a verifiable evidence bundle:
    boruna workflow run examples/compliance/soc2_audit_workflow --policy allow-all --record
    boruna evidence verify <bundle-dir>
    
  4. The verified bundle is your compliance artifact

Customisation

Each template README describes which steps to modify for your environment. Real integrations typically replace the gather_* or receive_* first step with a capability call to fetch live data.

std-authz

Role and permission enforcement via policies

Package: std.authz Version: 0.1.0 Capabilities required: none

Overview

std-authz provides a simple, policy-driven authorization model based on named roles and numeric privilege levels. Call authz_check in your update handler before applying any state change that requires a permission gate. Because all functions are pure and deterministic, authorization decisions are fully auditable and replayable from the evidence bundle.

Installation

Add to your package.ax.json dependencies:

"std.authz": "0.1.0"

API Reference

Types

AuthzPolicy

type AuthzPolicy {
    admin_role: String,
    editor_role: String,
    viewer_role: String,
    admin_level: Int,
    editor_level: Int,
    viewer_level: Int
}

Maps role names to numeric privilege levels. Higher levels have broader access.

Role

type Role { name: String, level: Int }

Permission

type Permission { resource: String, action: String }

RolePermission

type RolePermission { role_name: String, resource: String, action: String }

AuthzResult

type AuthzResult { allowed: Int, reason: String }
  • allowed — 1 if the operation is permitted, 0 if denied
  • reason — human-readable explanation for logging

Functions

authz_default_policy() -> AuthzPolicy

Returns the built-in three-tier policy: admin (level 100), editor (level 50), viewer (level 10).

authz_check(policy: AuthzPolicy, role_name: String, resource: String, action: String) -> AuthzResult

The primary authorization gate. Checks whether role_name is permitted to perform action on resource.

Supported actions: "read" (requires viewer+), "write" (requires editor+), "delete" (requires admin).

Parameters

  • policy — the active AuthzPolicy
  • role_name — the principal’s role string
  • resource — the resource being accessed (informational; used for the reason message)
  • action — "read", "write", or "delete"

Returns — AuthzResult with allowed: 1 or allowed: 0 and a descriptive reason.

Example

fn main() -> Int {
  let policy: AuthzPolicy = authz_default_policy()
  let result: AuthzResult = authz_check(policy, "editor", "documents", "write")
  result.allowed
}
authz_require_role(policy: AuthzPolicy, role_name: String, min_level: Int) -> AuthzResult

Returns allowed if the role’s numeric level meets or exceeds min_level. Useful for custom thresholds beyond the three built-in actions.

authz_role_level(policy: AuthzPolicy, role_name: String) -> Int

Returns the numeric privilege level for a role name. Returns 0 for unknown roles.

authz_guard_update(policy: AuthzPolicy, role_name: String, resource: String, action: String) -> Int

Convenience wrapper that calls authz_check and returns only the allowed integer — handy in if guards.

authz_is_admin(policy: AuthzPolicy, role_name: String) -> Int

Returns 1 if role_name matches the policy’s admin role.

authz_can_write(policy: AuthzPolicy, role_name: String) -> Int

Returns 1 if the role has editor-level or higher privileges.

authz_can_read(policy: AuthzPolicy, role_name: String) -> Int

Returns 1 if the role has viewer-level or higher privileges.

Capabilities

None. All functions are pure predicates with no side effects.

Notes / Limitations

  • The policy is a fixed three-role structure. Custom roles beyond admin/editor/viewer can be checked via authz_require_role with an explicit min_level.
  • Role names are compared by string equality; casing matters.
  • resource is passed through to the reason string but is not used in access-control logic — access depends solely on the action and role level.

std-db

Typed database query helpers

Package: std.db Version: 0.1.0 Capabilities required: db.query

Overview

std-db provides a builder API for constructing typed database queries and converting them to Effect values for execution. Build a Query with the CRUD helpers, refine it with db_where, db_order, and db_paginate, then call db_to_effect to hand it off to the runtime. The library also includes a Pagination type and helpers for navigating paged result sets.

Installation

Add to your package.ax.json dependencies:

"std.db": "0.1.0"

Your policy must grant db.query.

API Reference

Types

Effect

type Effect { kind: String, payload: String, callback_tag: String }

Query

type Query {
    operation: String,
    table: String,
    columns: String,
    conditions: String,
    order_by: String,
    limit_val: Int,
    offset_val: Int
}

A portable query descriptor. columns is a comma-separated list; conditions is a raw filter expression string.

Pagination

type Pagination { page: Int, per_page: Int, total: Int, total_pages: Int }

Computed pagination metadata. Use with pagination_has_next / pagination_has_prev in your view.

Functions

Query builders

db_select(table: String, columns: String) -> Query

Creates a SELECT query for the specified columns.

Example

fn main() -> Int {
  let q: Query = db_select("users", "id, name, email")
  let q2: Query = db_where(q, "active = 1")
  let q3: Query = db_order(q2, "name")
  let q4: Query = db_paginate(q3, 1, 20)
  let eff: Effect = db_to_effect(q4, "users_loaded")
  0
}
db_insert(table: String, columns: String, values: String) -> Query

Creates an INSERT query. columns and values are comma-separated strings corresponding to each other positionally.

db_update(table: String, sets: String, conditions: String) -> Query

Creates an UPDATE query. sets is a comma-separated list of column = value assignments.

db_delete(table: String, conditions: String) -> Query

Creates a DELETE query filtered by conditions.

Query modifiers

db_where(query: Query, condition: String) -> Query

Replaces the query’s condition expression. Call after the initial builder.

db_order(query: Query, column: String) -> Query

Sets the ORDER BY column.

db_limit(query: Query, limit_val: Int) -> Query

Sets a row limit.

db_offset(query: Query, offset_val: Int) -> Query

Sets a row offset.

db_paginate(query: Query, page: Int, per_page: Int) -> Query

Sets limit_val and offset_val from a 1-based page number and page size. Equivalent to calling db_limit and db_offset with the computed values.

Effect dispatch

db_to_effect(query: Query, callback_tag: String) -> Effect

Converts a Query to an Effect with kind: "db_query". Return this from update to execute the query.

Pagination helpers

pagination_info(page: Int, per_page: Int, total: Int) -> Pagination

Computes the Pagination record from a known total row count.

pagination_has_next(p: Pagination) -> Int

Returns 1 if there is a next page.

pagination_has_prev(p: Pagination) -> Int

Returns 1 if there is a previous page.

pagination_next_page(p: Pagination) -> Int

Returns the next page number, or the current page if already at the last.

pagination_prev_page(p: Pagination) -> Int

Returns the previous page number, or 1 if already at the first.

Capabilities

Requires db.query. Queries are executed by the runtime’s capability handler; the library itself produces only data structures.

Notes / Limitations

  • Condition strings (conditions, sets) are passed through verbatim to the runtime. The library does not perform sanitization; callers are responsible for safe parameterization.
  • db_paginate uses 1-based page numbers. Page 0 or negative values produce a negative offset.
  • total_pages is computed as ceil(total / per_page); when total == 0 the result is 1.

std-forms

Model-driven form engine with validation and state tracking

Package: std.forms Version: 0.1.0 Capabilities required: none

Overview

std-forms manages the lifecycle of multi-field forms inside framework apps. It provides two complementary state types — FieldState for individual inputs and FormState for the whole form — and a set of pure functions to transition between them. Use it in your update(State, Msg) -> UpdateResult handler to track dirty/touched flags, field errors, and submission state without writing that boilerplate yourself.

std-forms depends on std.validation for composable rule checking.

Installation

Add to your package.ax.json dependencies:

"std.forms": "0.1.0"

std.validation is pulled in automatically as a transitive dependency.

API Reference

Types

FieldState

type FieldState { name: String, value: String, touched: Int, dirty: Int, error: String }
  • touched — 1 if the user has focused the field at least once
  • dirty — 1 if the value has changed from its initial state
  • error — current validation error message, or "" when valid

FormState

type FormState {
    field_count: Int,
    submitted: Int,
    valid: Int,
    current_field: String,
    current_value: String,
    current_error: String
}
  • submitted — 1 after a successful form_submit
  • valid — 0 as soon as any field error is set; reset by form_clear_errors

Functions

Form lifecycle

form_init(field_count: Int) -> FormState

Creates a fresh form. Pass the number of fields to pre-declare capacity.

form_set_field(form: FormState, name: String, value: String) -> FormState

Records the most-recently changed field name and value. Call this on every input change message.

form_set_error(form: FormState, field: String, error: String) -> FormState

Attaches a validation error to a named field and marks the form invalid.

form_clear_errors(form: FormState) -> FormState

Clears all errors and resets valid to 1.

form_submit(form: FormState) -> FormState

Transitions the form to the submitted state if valid == 1. No-ops if the form is invalid.

form_reset(form: FormState) -> FormState

Resets all fields to their initial state while preserving field_count.

Form predicates

form_is_valid(form: FormState) -> Int

Returns 1 if no errors are set.

form_is_submitted(form: FormState) -> Int

Returns 1 after a successful submission.

Field lifecycle

field_init(name: String) -> FieldState

Creates a fresh field with an empty value and no errors.

field_set_value(field: FieldState, value: String) -> FieldState

Updates the field value and marks it dirty.

field_touch(field: FieldState) -> FieldState

Marks the field as touched (user has interacted with it).

field_set_error(field: FieldState, error: String) -> FieldState

Attaches a validation error message.

field_clear_error(field: FieldState) -> FieldState

Clears any error message.

field_is_valid(field: FieldState) -> Int

Returns 1 if the field has no error.

field_reset(field: FieldState) -> FieldState

Resets to empty, untouched, and error-free while keeping the field name.

Example

fn main() -> Int {
  let form: FormState = form_init(2)
  let name_field: FieldState = field_init("name")
  let name_typed: FieldState = field_set_value(name_field, "Alice")
  let name_visited: FieldState = field_touch(name_typed)
  let form2: FormState = form_set_field(form, "name", "Alice")
  let result: FormState = form_submit(form2)
  result.submitted
}

Capabilities

None. All state transitions are pure functions.

Notes / Limitations

  • FormState stores only the most-recently changed field. Apps with multiple fields should maintain per-field FieldState values in their own State record and use FormState for overall validity and submission tracking.
  • field_count is informational only; the engine does not iterate over fields.

std-http

Safe typed HTTP effect wrappers

Package: std.http Version: 0.1.0 Capabilities required: net.fetch

Overview

std-http wraps outbound HTTP calls as typed Effect values. Your update handler returns these effects; the Boruna runtime dispatches them through the net.fetch capability and delivers responses back via the named callback_tag. This keeps network I/O out of pure logic and makes every request auditable in the evidence bundle. The library also provides retry configuration helpers and status-code predicates.

Installation

Add to your package.ax.json dependencies:

"std.http": "0.1.0"

Your workflow or app policy must grant net.fetch to the step that uses this library.

API Reference

Types

Effect

type Effect { kind: String, payload: String, callback_tag: String }

Returned by all request-building functions. Pass it back from update to trigger the HTTP call.

HttpRequest

type HttpRequest { method: String, url: String, body: String, content_type: String }

A fully described request for use with http_request.

HttpResponse

type HttpResponse { status: Int, body: String }

Shape of the response delivered to the callback_tag handler.

RetryConfig

type RetryConfig { max_retries: Int, backoff_ms: Int, multiplier: Int }

Exponential backoff parameters used with http_next_backoff.

Functions

Request builders

http_get(url: String, callback_tag: String) -> Effect

Produces a GET request effect.

Example

fn main() -> Int {
  let eff: Effect = http_get("https://api.example.com/items", "items_loaded")
  0
}
http_post(url: String, body: String, callback_tag: String) -> Effect

Produces a POST request effect with the given body.

http_put(url: String, body: String, callback_tag: String) -> Effect

Produces a PUT request effect.

http_delete(url: String, callback_tag: String) -> Effect

Produces a DELETE request effect.

http_request(req: HttpRequest, callback_tag: String) -> Effect

Produces an effect from a fully specified HttpRequest — use when you need to set content_type or other fields explicitly.

Retry helpers

http_default_retry() -> RetryConfig

Returns a sensible default: 3 retries, 1000 ms base backoff, multiplier 2.

http_retry_config(max_retries: Int, backoff_ms: Int, multiplier: Int) -> RetryConfig

Constructs a custom retry configuration.

http_next_backoff(config: RetryConfig, attempt: Int) -> Int

Returns the wait duration in milliseconds before the next retry for a given attempt index (0-based). Uses fixed exponential steps up to attempt == 2; subsequent attempts cap at backoff_ms * multiplier^2.

Status-code predicates

http_is_success(status: Int) -> Int

Returns 1 for 2xx status codes.

http_is_error(status: Int) -> Int

Returns 1 for 4xx and 5xx status codes.

http_should_retry(status: Int) -> Int

Returns 1 for status 429 (Too Many Requests) or any 5xx — the conditions under which retrying is appropriate.

http_parse_status(payload: String) -> Int

Parses a status code string to Int. Returns 0 if parsing fails.

Capabilities

Requires net.fetch. The VM’s CapabilityGateway enforces this at runtime; the call is rejected if the active policy does not include net.fetch.

Notes / Limitations

  • All functions produce Effect values — they do not perform any I/O themselves. Actual network calls happen in the runtime after the update function returns.
  • http_next_backoff implements three-step exponential backoff only. If attempt >= 3 the backoff is not further increased in the current implementation.
  • http_parse_status is a stub that returns 0; the runtime is expected to set the status field directly on the response record delivered to the callback.

std-json

JSON-flavoured data extraction and manipulation utilities

Package: std.json Version: 0.1.0 Capabilities required: none

Overview

std-json provides pure-functional helpers for building JSON-formatted strings from typed Boruna values. It covers the four JSON scalar types (string, number, boolean, null), object wrapping, and array wrapping. Because it has no capability requirements it can be used in any step regardless of policy. All functions are deterministic and produce no side effects.

Installation

Add to your package.ax.json dependencies:

"std.json": "0.1.0"

No capability grants are required.

API Reference

Types

JsonResult

type JsonResult { ok: Bool, value: String, error: String }

A tagged result for operations that may fail. When ok is true, value holds the result; when false, error holds the reason.

Functions

Result constructors

json_ok(value: String) -> JsonResult

Wraps a successful string result in a JsonResult with ok = true.

json_err(error: String) -> JsonResult

Wraps an error message in a JsonResult with ok = false.

Field serialisers

json_string_field(key: String, value: String) -> String

Returns a JSON key-value fragment for a string value: "key": "value".

Example

fn main() -> Int {
  let field: String = json_string_field("name", "Alice")
  // field == "\"name\": \"Alice\""
  let obj: String = json_object(field)
  // obj == "{\"name\": \"Alice\"}"
  0
}
json_int_field(key: String, value: Int) -> String

Returns a JSON key-value fragment for an integer value: "key": <n>.

json_bool_field(key: String, value: Bool) -> String

Returns a JSON key-value fragment for a boolean value: "key": true or "key": false.

json_null_field(key: String) -> String

Returns a JSON key-value fragment for a null value: "key": null.

Container builders

json_object(fields: String) -> String

Wraps a pre-serialised fields string in {...}. Combine multiple fields with string concatenation before passing.

json_array_wrap(items: String) -> String

Wraps a pre-serialised items string in [...].

Utilities

json_escape(s: String) -> String

Escapes a string for safe inclusion in a JSON value. Correctly escapes " and \ characters.

int_to_string(n: Int) -> String

Converts an integer to its decimal string representation. Wraps __builtin_int_to_string.

Capabilities

None. std-json is pure-functional with no side effects.

Notes

  • There is no recursive structure support; callers must build nested JSON by concatenating field strings manually before passing to json_object.

Version History

VersionChange
0.1.0Initial release. JsonResult type; json_ok, json_err, json_string_field, json_int_field, json_bool_field, json_null_field, json_object, json_array_wrap, json_escape, int_to_string.

std-llm

Typed LLM effect wrappers for prompting and structured generation

Package: std.llm Version: 0.1.0 Capabilities required: llm.call

Overview

std-llm wraps LLM calls as typed Effect values. Your update handler returns these effects; the Boruna runtime dispatches them through the llm.call capability and delivers responses back via the named callback_tag. This keeps model I/O out of pure logic and makes every prompt auditable in the evidence bundle. Temperature is stored as Int * 100 (e.g. 72 = 0.72) to avoid floating-point precision issues.

Installation

Add to your package.ax.json dependencies:

"std.llm": "0.1.0"

Your workflow or app policy must grant llm.call to the step that uses this library.

API Reference

Types

Effect

type Effect { kind: String, payload: String, callback_tag: String }

Returned by all request-building functions. Pass it back from update to trigger the LLM call. kind is either "llm_call" or "llm_json_call".

LlmRequest

type LlmRequest { system_prompt: String, user_prompt: String, max_tokens: Int, temperature: Int }

A fully described request for use with llm_call or llm_json_call. temperature is Int * 100 (e.g. 72 = 0.72).

LlmResponse

type LlmResponse { content: String, tokens_used: Int, finish_reason: String }

Shape of the response delivered to the callback_tag handler.

Functions

llm_prompt(system: String, user: String, callback_tag: String) -> Effect

Convenience function: builds a default LlmRequest (1024 max tokens, temperature 0.72) and emits an "llm_call" effect.

Example

fn main() -> Int {
  let eff: Effect = llm_prompt(
    "You are a helpful assistant.",
    "Summarize this document.",
    "summary_done"
  )
  0
}

llm_call(req: LlmRequest, callback_tag: String) -> Effect

Emits an "llm_call" effect from a fully specified LlmRequest — use when you need to control max_tokens or temperature explicitly.

llm_json_call(req: LlmRequest, callback_tag: String) -> Effect

Emits an "llm_json_call" effect, signalling to the runtime that the model should be prompted for structured JSON output.

default_llm_request(system: String, user: String) -> LlmRequest

Constructs an LlmRequest with default values: max_tokens = 1024, temperature = 72. Use when you want to inspect or mutate the request before passing it to llm_call.

Capabilities

Requires llm.call. The VM’s CapabilityGateway enforces this at runtime; the call is rejected if the active policy does not include llm.call.

Notes / Limitations

  • All functions produce Effect values — they do not perform any I/O themselves. Actual model calls happen in the runtime after the update function returns.
  • temperature is encoded as Int * 100; 72 means 0.72. There is no float conversion in-language.
  • The payload field of the produced Effect is system_prompt ++ "|" ++ user_prompt; the runtime splits on | to reconstruct the two prompts.

Version History

VersionChange
0.1.0Initial release. LlmRequest, LlmResponse, Effect types; llm_call, llm_prompt, llm_json_call, default_llm_request functions.

std-notifications

Notification queue and toast helpers

Package: std.notifications Version: 0.1.0 Capabilities required: time.now

Overview

std-notifications manages a bounded in-memory queue of toast-style notifications. Add messages with the level-specific helpers (notification_success, notification_error, etc.), dismiss them by ID, and schedule auto-dismiss timers with notification_dismiss_effect. All queue management is pure; only auto-dismiss produces a timer Effect.

Installation

Add to your package.ax.json dependencies:

"std.notifications": "0.1.0"

Your policy must grant time.now if you use notification_dismiss_effect.

API Reference

Types

Effect

type Effect { kind: String, payload: String, callback_tag: String }

Notification

type Notification { id: Int, level: String, message: String, dismiss_ms: Int }
  • id — unique monotonic integer assigned at push time
  • level — "info", "success", "warning", or "error"
  • dismiss_ms — milliseconds until auto-dismiss; 0 means no auto-dismiss

NotificationQueue

type NotificationQueue {
    next_id: Int,
    count: Int,
    max_visible: Int,
    last_level: String,
    last_message: String,
    last_id: Int
}
  • next_id — the ID that will be assigned to the next pushed notification
  • count — number of currently visible notifications
  • max_visible — cap beyond which new pushes should be blocked or the oldest dropped
  • last_level / last_message / last_id — fields of the most recently pushed notification

Functions

Initialization

notification_init(max_visible: Int) -> NotificationQueue

Creates an empty queue with the given visible-notification cap.

Pushing notifications

notification_info(queue: NotificationQueue, level: String, message: String) -> NotificationQueue

Pushes a notification with level "info" and dismiss_ms: 3000.

notification_success(queue: NotificationQueue, message: String) -> NotificationQueue

Pushes a success notification with dismiss_ms: 3000.

notification_warning(queue: NotificationQueue, message: String) -> NotificationQueue

Pushes a warning notification with dismiss_ms: 5000.

notification_error(queue: NotificationQueue, message: String) -> NotificationQueue

Pushes an error notification with dismiss_ms: 0 (sticky — does not auto-dismiss).

notification_push(queue: NotificationQueue, level: String, message: String, dismiss_ms: Int) -> NotificationQueue

Low-level push with explicit level and dismiss duration.

Example

fn main() -> Int {
  let q: NotificationQueue = notification_init(5)
  let q1: NotificationQueue = notification_success(q, "Changes saved")
  let q2: NotificationQueue = notification_error(q1, "Upload failed")
  q2.count
}

Constructing a notification record

notification_make(id: Int, level: String, message: String, dismiss_ms: Int) -> Notification

Constructs a Notification value directly. Use when rendering the queue in your view.

Dismissing

notification_dismiss(queue: NotificationQueue, id: Int) -> NotificationQueue

Decrements count by one. The id parameter is informational; the implementation does not track individual items in the current version.

notification_dismiss_effect(id: Int, delay_ms: Int) -> Effect

Produces a timer effect that fires after delay_ms milliseconds. Wire the "notification_dismissed" callback to call notification_dismiss in your update handler.

Predicates

notification_count(queue: NotificationQueue) -> Int

Returns the current number of visible notifications.

notification_is_full(queue: NotificationQueue) -> Int

Returns 1 if count >= max_visible.

Capabilities

time.now is required for the timer effect produced by notification_dismiss_effect. Queue-management functions are pure and need no capability.

Notes / Limitations

  • The queue does not store individual Notification records internally. last_id, last_level, and last_message reflect only the most recently pushed item. To render all visible notifications your app state should maintain a list of Notification values alongside the NotificationQueue.
  • notification_dismiss decrements the count without validating the id. Calling it more times than there are visible notifications will clamp at 0.
  • Auto-dismiss (dismiss_ms > 0) requires your update handler to return the timer effect and handle the "notification_dismissed" callback.

std-routing

Declarative routing model

Package: std.routing Version: 0.1.0 Capabilities required: none

Overview

std-routing gives framework apps a lightweight, deterministic routing layer. Routes are declared as Route records, and matching is done with pure functions — no parsing magic, no global history object. Define your routes once, call route_match_first in your update handler when the path changes, and use route_is_active in your view to highlight the current link.

Installation

Add to your package.ax.json dependencies:

"std.routing": "0.1.0"

API Reference

Types

Route

type Route { name: String, path: String, param_count: Int }
  • name — a stable identifier used throughout the app (e.g. "home", "users")
  • path — the URL path string to match against (e.g. "/", "/users")
  • param_count — number of dynamic parameters in the path (0 for static routes)

RouteMatch

type RouteMatch { matched: Int, route_name: String, param1: String, param2: String }
  • matched — 1 if a route matched, 0 otherwise
  • route_name — the name of the matched route, "" on no match
  • param1 / param2 — extracted path parameters (up to two)

Functions

route_define(name: String, path: String) -> Route

Declares a static route with no dynamic parameters.

Example

fn main() -> Int {
  let home: Route = route_define("home", "/")
  let users: Route = route_define("users", "/users")
  let settings: Route = route_define("settings", "/settings")
  let m: RouteMatch = route_match_first(home, users, settings, "/users")
  m.matched
}
route_define_with_param(name: String, path: String, param_count: Int) -> Route

Declares a route that expects param_count dynamic path segments.

route_match_path(route: Route, path: String) -> RouteMatch

Checks a single route against path. Returns a no-match result if route.path != path.

route_match_first(r1: Route, r2: Route, r3: Route, path: String) -> RouteMatch

Tries r1, r2, and r3 in order and returns the first match. Returns the result of r3 (which may be a no-match) if none of the first two matched. Use as a three-entry router table.

route_no_match() -> RouteMatch

Returns a zero-value no-match result. Use as a fallback or default.

route_navigate(route_name: String) -> String

Returns the route name as a navigation token. The framework dispatches this to the runtime to push the new path.

route_navigate_with_param(route_name: String, param: String) -> String

Like route_navigate but carries a single parameter.

route_is_active(current_route: String, route_name: String) -> Int

Returns 1 if current_route == route_name. Use in view to highlight active navigation links.

Capabilities

None. All functions are pure.

Notes / Limitations

  • route_match_path uses exact string equality. Dynamic path segments (e.g. /users/42) require the calling app to extract parameters before matching, or use param_count as a hint to implement custom segment splitting.
  • route_match_first handles exactly three routes. For larger route tables, nest calls or use a series of route_match_path checks in your update handler.
  • route_navigate returns the route name string as a navigation token; the actual URL push is performed by the runtime, not the library.

std-storage

Typed local persistence abstraction

Package: std.storage Version: 0.1.0 Capabilities required: fs.read, fs.write

Overview

std-storage provides a namespaced key-value persistence layer as Effect values. Operations are described purely at the .ax level; actual I/O happens in the Boruna runtime. Entries carry a monotonic version counter which makes conflict detection straightforward and keeps all storage operations fully replay-compatible.

Installation

Add to your package.ax.json dependencies:

"std.storage": "0.1.0"

Your policy must grant both fs.read and fs.write.

API Reference

Types

Effect

type Effect { kind: String, payload: String, callback_tag: String }

StorageKey

type StorageKey { namespace: String, key: String }

A typed composite key. The namespace scopes keys to a logical bucket (e.g. "app", "cache").

StorageEntry

type StorageEntry { namespace: String, key: String, value: String, version: Int }

A versioned key-value record. value is stored as a JSON string by convention. version is a monotonic integer incremented by storage_bump_version.

Functions

Key construction

storage_key(namespace: String, key: String) -> StorageKey

Constructs a typed key. Useful for passing keys around without losing the namespace.

Effect builders

storage_get(namespace: String, key: String, callback_tag: String) -> Effect

Produces a read effect. The runtime delivers the stored value to callback_tag.

Example

fn main() -> Int {
  let eff: Effect = storage_get("app", "settings", "settings_loaded")
  0
}
storage_set(namespace: String, key: String, value: String, callback_tag: String) -> Effect

Produces a write effect. value is written at the given key.

storage_delete(namespace: String, key: String, callback_tag: String) -> Effect

Produces a delete effect. The key is removed from the namespace.

storage_list(namespace: String, callback_tag: String) -> Effect

Produces a list effect. The runtime delivers all keys in the namespace to callback_tag.

Entry helpers

storage_make_entry(namespace: String, key: String, value: String, version: Int) -> StorageEntry

Constructs a StorageEntry with an explicit version. Use when deserializing a stored entry.

storage_bump_version(entry: StorageEntry, new_value: String) -> StorageEntry

Returns a new entry with value updated and version incremented by one.

Example: optimistic update

fn main() -> Int {
  let entry: StorageEntry = storage_make_entry("app", "counter", "0", 1)
  let updated: StorageEntry = storage_bump_version(entry, "1")
  updated.version
}
storage_is_newer(a: StorageEntry, b: StorageEntry) -> Int

Returns 1 if a.version > b.version. Use for last-write-wins conflict resolution.

Capabilities

Requires fs.read for read/list operations and fs.write for write/delete operations. The capabilities are enforced separately — a step that only reads can be granted fs.read alone.

Notes / Limitations

  • value is an untyped String. By convention store JSON, but the library does not enforce any encoding.
  • Versioning is local only; it does not synchronize with remote storage. For distributed conflict resolution use std-sync.
  • All four effect kinds (storage_read, storage_write, storage_delete, storage_list) are replay-safe: the runtime records the result in the EventLog and replays it deterministically.

std-sync

Offline queue and conflict resolution helpers

Package: std.sync Version: 0.1.0 Capabilities required: net.fetch

Overview

std-sync tracks local-vs-remote divergence and manages an offline edit queue inside framework apps. All conflict-resolution logic is pure; only sync_effect produces a network Effect. Use it to build apps that continue working offline and flush pending changes when connectivity is restored.

Installation

Add to your package.ax.json dependencies:

"std.sync": "0.1.0"

Your policy must grant net.fetch for sync_effect calls to succeed.

API Reference

Types

Effect

type Effect { kind: String, payload: String, callback_tag: String }

SyncState

type SyncState {
    online: Int,
    pending_count: Int,
    synced_count: Int,
    status: String,
    local_version: Int,
    remote_version: Int,
    conflicts: Int,
    resolved: Int
}
  • online — 1 when connected, 0 when offline
  • pending_count — number of local edits not yet synced
  • synced_count — cumulative count of successfully synced operations
  • status — one of "idle", "queued", "syncing", "synced", "offline", "conflict", "resolving"
  • local_version / remote_version — monotonic counters; divergence indicates pending changes
  • conflicts / resolved — total conflicts detected vs resolved

Functions

Initialization

sync_init() -> SyncState

Returns a fresh sync state: online, no pending edits, version 0.

State transitions

sync_queue_edit(state: SyncState) -> SyncState

Records one new local edit. Increments pending_count and local_version. Sets status to "syncing" if online, "queued" if offline.

sync_mark_synced(state: SyncState) -> SyncState

Acknowledges one successfully synced edit. Decrements pending_count and advances remote_version to match local_version. Status becomes "synced" when the queue is drained.

sync_go_offline(state: SyncState) -> SyncState

Marks the state as offline and sets status to "offline".

sync_go_online(state: SyncState) -> SyncState

Marks the state as online. Status becomes "syncing" if there are pending edits, "synced" otherwise.

sync_detect_conflict(state: SyncState) -> SyncState

Increments the conflict counter and sets status to "conflict". Call when the server rejects an edit due to a version mismatch.

sync_resolve_conflict(state: SyncState, strategy: String) -> SyncState

Increments resolved, bumps local_version, and sets status to "resolving". strategy is passed for audit purposes; conflict resolution logic is app-specific.

Predicates

sync_needs_push(state: SyncState) -> Int

Returns 1 if there are pending edits and the device is online — the signal to fire a sync effect.

sync_is_idle(state: SyncState) -> Int

Returns 1 if status is "idle" or "synced".

sync_has_conflicts(state: SyncState) -> Int

Returns 1 if any unresolved conflicts remain.

Effects

sync_effect(endpoint: String, callback_tag: String) -> Effect

Produces an HTTP request effect targeting endpoint. Return this from update when sync_needs_push is true.

Example

fn main() -> Int {
  let s0: SyncState = sync_init()
  let s1: SyncState = sync_queue_edit(s0)
  let s2: SyncState = sync_go_offline(s1)
  let s3: SyncState = sync_queue_edit(s2)
  let s4: SyncState = sync_go_online(s3)
  s4.pending_count
}

Capabilities

Requires net.fetch for sync_effect. All state-transition functions are pure and require no capability.

Notes / Limitations

  • sync_resolve_conflict records a resolution but does not merge data. The calling app is responsible for choosing and applying the correct value before calling this function.
  • sync_needs_push checks connectivity and pending count only; rate-limiting and retry backoff are not built in. Combine with std-http’s RetryConfig for production use.

std-testing

High-level test helpers for framework apps

Package: std.testing Version: 0.1.0 Capabilities required: none

Overview

std-testing provides assertion functions and aggregation helpers for writing deterministic unit tests directly in .ax. Use it inside fn main() test programs or alongside the boruna framework test command to verify app behavior. Because everything is pure, test programs are fully deterministic and replay-safe.

Installation

Add to your package.ax.json dependencies:

"std.testing": "0.1.0"

API Reference

Types

TestResult

type TestResult { passed: Int, label: String, detail: String }
  • passed — 1 if the assertion succeeded, 0 if it failed
  • label — the name of this test case
  • detail — "ok" on success, or a short failure reason

TestSummary

type TestSummary { total: Int, passed: Int, failed: Int }

Aggregate result for a group of tests.

Functions

Assertions

assert_eq_int(actual: Int, expected: Int, label: String) -> TestResult

Passes if actual == expected.

Example

fn main() -> Int {
  let r: TestResult = assert_eq_int(2 + 2, 4, "addition")
  r.passed
}
assert_eq_string(actual: String, expected: String, label: String) -> TestResult

Passes if actual == expected.

assert_true(value: Int, label: String) -> TestResult

Passes if value == 1.

assert_false(value: Int, label: String) -> TestResult

Passes if value == 0.

assert_gt(actual: Int, threshold: Int, label: String) -> TestResult

Passes if actual > threshold.

assert_lt(actual: Int, threshold: Int, label: String) -> TestResult

Passes if actual < threshold.

assert_not_empty(value: String, label: String) -> TestResult

Passes if value != "".

Aggregation

test_summary(t1: TestResult, t2: TestResult, t3: TestResult) -> TestSummary

Computes a summary across exactly three test results.

test_all_passed_2(t1: TestResult, t2: TestResult) -> Int

Returns 1 if both results passed.

test_all_passed_3(t1: TestResult, t2: TestResult, t3: TestResult) -> Int

Returns 1 if all three results passed.

Example: full test program

fn main() -> Int {
  let r1: TestResult = assert_eq_int(1 + 1, 2, "basic addition")
  let r2: TestResult = assert_true(1, "literal true")
  let r3: TestResult = assert_not_empty("hello", "non-empty string")
  let summary: TestSummary = test_summary(r1, r2, r3)
  summary.failed
}

A main that returns 0 indicates all tests passed (zero failures).

Capabilities

None. All functions are pure.

Notes / Limitations

  • test_summary handles exactly three tests. To summarize more, compute intermediate summaries or use test_all_passed_2 / test_all_passed_3 in a chain.
  • There is no test runner built into the library itself. Run test programs with boruna run <file>.ax or boruna framework test for framework apps.
  • Failure detail is limited to "mismatch", "expected true", "expected false", "not greater than", "not less than", or "was empty". Custom detail messages are not yet supported; use label to identify the assertion.

std-ui

Declarative UI primitives for framework apps

Package: std.ui Version: 0.1.0 Capabilities required: none

Overview

std-ui provides a small set of pure functions that return UINode values — the building blocks of a Boruna framework app’s view layer. Every function is deterministic and side-effect-free. Wire these together in your view(State) -> UINode implementation to compose layouts, controls, and data-display elements.

Installation

Add to your package.ax.json dependencies:

"std.ui": "0.1.0"

API Reference

Types

UINode

type UINode { tag: String, text: String }

The universal tree node returned by every std-ui function. The tag identifies the element kind; text carries the primary display string. The framework runtime interprets these fields when rendering.

Functions

Layout

row(child1: UINode, child2: UINode) -> UINode

Places two nodes side-by-side in a horizontal row.

column(child1: UINode, child2: UINode) -> UINode

Stacks two nodes vertically.

stack(child1: UINode, child2: UINode) -> UINode

Overlays two nodes in a Z-stack (front-to-back).

Containers

container(content: UINode) -> UINode

Wraps a node in a generic container — useful for spacing or styling boundaries.

card(title: String, content: UINode) -> UINode

Renders a titled card surface around content.

Parameters

  • title — display title shown at the top of the card
  • content — the inner UINode to render inside
section(heading: String, content: UINode) -> UINode

A named section with a visible heading above content.

Controls

button(label: String, on_click: String) -> UINode

An interactive button.

Parameters

  • label — text displayed on the button
  • on_click — message tag dispatched when clicked

Example

fn main() -> Int {
  let btn: UINode = button("Save", "save_clicked")
  0
}
input(name: String, value: String, on_change: String) -> UINode

A text input field.

Parameters

  • name — field identifier
  • value — current value to display
  • on_change — message tag dispatched on every change
select_field(name: String, selected: String, on_change: String) -> UINode

A dropdown select field.

checkbox(name: String, checked: Int, on_change: String) -> UINode

A checkbox. checked: 1 renders the box checked; 0 renders it unchecked.

Data display

table_view(headers: String, row_count: Int) -> UINode

A tabular view. headers is a comma-separated list of column names. row_count indicates how many rows the runtime should render.

error_display(message: String) -> UINode

Shows an error message in a styled error block.

text(content: String) -> UINode

Renders a plain text node.

badge(label: String, variant: String) -> UINode

A small inline badge. variant hints at the visual style (e.g. "success", "warning", "error").

Helpers

empty_node() -> UINode

A no-op placeholder node that renders nothing. Use when a conditional branch needs to return a UINode without visible output.

divider() -> UINode

A horizontal rule / visual separator.

Capabilities

None. All functions are pure and return data structures only — no I/O occurs during view construction.

Notes / Limitations

  • UINode currently carries only tag and text. Attributes such as CSS classes, event payloads, and children beyond the first two are not yet representable in the type. This is a v0.x constraint; the type will be expanded in a future release.
  • row, column, and stack each accept exactly two children. Composing more requires nesting calls.
  • string_length and other built-ins used internally are resolved by the Boruna runtime.

std-validation

Reusable composable validation rules

Package: std.validation Version: 0.1.0 Capabilities required: none

Overview

std-validation provides a small set of pure, composable validation functions. Each returns a ValidationResult that carries a pass/fail flag and an error message. Chain results with validation_merge to apply multiple rules in sequence and surface the first failure. std-forms depends on this library for field-level validation.

Installation

Add to your package.ax.json dependencies:

"std.validation": "0.1.0"

API Reference

Types

ValidationError

type ValidationError { field: String, message: String }

ValidationResult

type ValidationResult { valid: Int, error_field: String, error_message: String }
  • valid — 1 if the rule passed, 0 if it failed
  • error_field — the field name on failure, "" on success
  • error_message — a short description of what failed, "" on success

Functions

String rules

validate_required(field: String, value: String) -> ValidationResult

Fails if value is an empty string.

Example

fn main() -> Int {
  let r: ValidationResult = validate_required("email", "")
  r.valid
}
validate_not_empty(field: String, value: String) -> ValidationResult

Alias for validate_required.

validate_min_length(field: String, value: String, min: Int) -> ValidationResult

Fails if the string length is less than min.

validate_max_length(field: String, value: String, max: Int) -> ValidationResult

Fails if the string length exceeds max.

validate_numeric(field: String, value: String) -> ValidationResult

Fails if value cannot be parsed as an integer.

validate_equals(field: String, actual: String, expected: String) -> ValidationResult

Fails if actual != expected. Use for password confirmation fields.

Integer rules

validate_min_value(field: String, value: Int, min: Int) -> ValidationResult

Fails if value < min.

validate_max_value(field: String, value: Int, max: Int) -> ValidationResult

Fails if value > max.

Composition

validation_ok() -> ValidationResult

Returns a pre-built passing result. Use as the initial accumulator in a chain.

validation_fail(field: String, message: String) -> ValidationResult

Returns a pre-built failing result with a custom message.

validation_merge(a: ValidationResult, b: ValidationResult) -> ValidationResult

Returns a if it failed, otherwise returns b. Use to apply rules in priority order and surface the first failure.

Example: chaining multiple rules

fn main() -> Int {
  let r1: ValidationResult = validate_required("username", "alice")
  let r2: ValidationResult = validate_min_length("username", "alice", 3)
  let r3: ValidationResult = validate_max_length("username", "alice", 20)
  let result: ValidationResult = validation_merge(validation_merge(r1, r2), r3)
  result.valid
}

Capabilities

None. All functions are pure.

Notes / Limitations

  • string_length is a runtime built-in; the stub in core.ax returns 0 and is replaced at compile time.
  • validate_numeric uses try_parse_int (a pattern-matching built-in); the library does not support float parsing.
  • validation_merge returns the first failure only — it does not accumulate multiple errors. If you need to collect all errors, maintain a list of ValidationResult in your app state.

Boruna Versioned Specifications

This directory holds formal, versioned specifications for the surfaces Boruna commits to keeping stable.

Each spec carries a language_version / format_version / schema_version field in its front matter or top-level shape. Implementations against a 1.x spec MUST keep working against any later 1.y (y >= x).

Current specs

SurfaceLatestStatusSprintReader constant
.ax language1.0stableW1-Bboruna_compiler::LANGUAGE_VERSION
Bytecode format1.0stableW9-Aboruna_bytecode::BYTECODE_VERSION
Evidence bundle format1.0stableW1-Cboruna_orchestrator::BUNDLE_FORMAT_VERSION
Workflow DAG schema1.0stableW4boruna_orchestrator::WORKFLOW_DAG_SCHEMA_VERSION

The narrative companion to the bytecode spec lives at docs/bytecode-spec.md; the formal spec at bytecode-1.0.md wins on any disagreement.

Authoring rules

  1. Specs are prescriptive, not descriptive. They are the authority. Reference docs (under docs/reference/) and concept docs (under docs/concepts/) are interpretive.
  2. Each spec MUST declare its version, status, and last-revised date in YAML front matter.
  3. Each spec MUST include a backwards-compatibility commitment for its current major line.
  4. Once a spec at version M.N is shipped in a release tag, it is frozen. Corrections that change behavior require bumping to M.(N+1) (additive) or (M+1).0 (breaking).
  5. Frozen specs MAY be edited only for clarifications that do not change observable conformance — typo fixes, wording, examples.

Versioning policy

MAJOR.MINOR decimal:

  • Major bump (1.0 → 2.0) — breaking change. A 1.x program may stop working.
  • Minor bump (1.0 → 1.1) — additive only. Every 1.0 program still works.

There is no patch version on specs; clarifying edits keep the same minor.

Reader contract

  • Hard reject across a major. A reader built for N.x MUST refuse N+1.0 documents with a typed Unsupported*Version error rather than guess.
  • Forward-compat within a major. A reader built for N.x MUST accept N.y documents (y >= x) and silently ignore unknown additive fields.
  • Replay invariant. Versions feed into the canonical-JSON serialization that produces workflow_hash/bundle hashes, binding evidence to a specific schema generation.

.ax language


spec_id: ax-language language_version: “1.0” status: stable last_revised: 2026-04-28 audience: language implementers, compiler authors, security auditors

.ax Language Specification — Version 1.1

This document is the formal specification of the .ax source language. It is the authoritative reference for any independent implementation of an .ax parser, type checker, or compiler.

The narrative, example-driven companion is docs/reference/ax-language.md. When the two disagree, this spec wins.

The reference implementation lives in crates/llmc/ (lexer, parser, typechecker, codegen) of the Boruna repository. Implementations against this spec MAY use a different code organization but MUST be observationally equivalent to the rules below.

1. Conformance and versioning

1.1 Version identifier

The current language version is 1.1. Implementations MUST expose this value programmatically. Version 1.1 is 1.0 plus the additions listed in §13; every 1.0 program is a 1.1 program.

In the reference implementation:

#![allow(unused)]
fn main() {
// crates/llmc/src/lib.rs
pub const LANGUAGE_VERSION: &str = "1.1";
}

The version string is a <major>.<minor> decimal number. A program written against 1.x MUST compile against any 1.y implementation where y >= x.

1.2 Backwards-compatibility commitment for 1.x

Within the 1.x line:

  1. Additive only. New keywords, types, opcodes, and capabilities MAY be added.
  2. No renames. The names Int, Float, String, Bool, Unit, Option, Result, List, Map, Some, None, Ok, Err, every reserved word in §2.4, and every capability in §6.2 are frozen for 1.x.
  3. No removed builtins. A function or operator that compiles in 1.x MUST continue to compile in any later 1.y.
  4. No tightened type rules. A program that type-checks in 1.x MUST type-check in 1.y >= 1.x. Type rules MAY be relaxed (more programs accepted) but MAY NOT be tightened.
  5. Capability annotation set is monotonic. A capability listed on a function signature in 1.x remains valid in 1.y >= 1.x. New capabilities added in later 1.y minor versions MUST be optional — programs that do not request them remain valid.
  6. Reserved words for future use. See §2.4. Implementations MUST tokenize them as reserved and MUST reject them outside of well-formed positions, even though they have no defined semantics in 1.0.

Breaking changes (renames, removals, tightening) are deferred to 2.0 or later.

1.3 Conformance levels

A 1.0 implementation MUST:

  • Accept every program admitted by the grammar in §3 that satisfies the type rules in §4.
  • Reject every program that violates the type rules in §4.
  • Enforce capability declarations (§6) at the boundary defined by the host runtime.
  • Preserve determinism (§7).

A 1.0 implementation MAY:

  • Emit additional diagnostics or hints.
  • Provide additional optimization passes, as long as they do not change observable behavior of well-formed programs.

2. Lexical structure

The source text is a UTF-8 byte string. Tokenization is greedy left-to-right, longest-match.

2.1 Whitespace and line endings

Whitespace characters are: ' ' (U+0020), '\t' (U+0009), '\r' (U+000D), '\n' (U+000A). Whitespace is significant only as a token separator. Statements are terminated by newlines (no semicolons).

2.2 Comments

LineComment ::= "//" {any character except newline} newline

Block comments are NOT defined in 1.0.

2.3 Identifiers

Identifier ::= IdStart IdContinue*
IdStart    ::= "a".."z" | "A".."Z" | "_"
IdContinue ::= IdStart | "0".."9"

Identifiers are case-sensitive. The identifier _ (single underscore) is a wildcard in patterns (§5) but is otherwise a normal identifier.

2.4 Keywords and reserved words

Keywords (have meaning; MUST NOT be used as identifiers):

fn   let   if   else   match   record   enum   true   false
Some None Ok  Err
mut  while for  in                                  (added in 1.1, §4.5)

Type names treated as reserved identifiers (§4.1):

Int Float String Bool Unit Option Result List Map

Reserved-for-future-use (a 1.0 implementation MUST tokenize these as reserved and reject them in identifier position; they have no semantics in 1.0):

loop return break continue trait impl import module
where as pub priv async await yield static const ref self Self
type spawn actor receive

(mut, for, while and in moved from this list to the keywords in 1.1. The reference implementation already gives some of the remaining words a meaning, for example return, import, spawn and receive; those are not yet specified here.)

2.5 Literals

IntLit    ::= ["-"] Digit+
FloatLit  ::= ["-"] Digit+ "." Digit+
BoolLit   ::= "true" | "false"
StringLit ::= "\"" {StringChar} "\""
StringChar::= any UTF-8 codepoint except "\"" and "\\", or one of the escapes:
              "\\\\" | "\\\"" | "\\n" | "\\t" | "\\r"
UnitLit   ::= "(" ")"
Digit     ::= "0".."9"

IntLit MUST fit in a 64-bit signed integer. FloatLit MUST be a finite IEEE 754 double-precision value.

2.6 Operators and punctuation

+  -  *  /  %      arithmetic
== != <  <= >  >=  comparison
&& || !            logical
=                  binding
->                 return type
=>                 match arm
::                 enum variant
..                 record spread
{ } ( ) [ ]        delimiters
, : .              separators / field access
!{ }               capability annotation opener / closer
"  \\              string delimiter / escape

3. Syntactic grammar (EBNF)

EBNF conventions: {X} = zero or more, [X] = optional, | = alternation, "x" = terminal, Newline = a newline outside any expression.

3.1 Programs

Program ::= {Item}
Item    ::= FnDecl | RecordDecl | EnumDecl

Every .ax source file is a Program. A standalone executable program MUST contain a top-level item fn main() -> Int.

3.2 Declarations

FnDecl     ::= "fn" Identifier "(" [Params] ")" "->" Type [CapAnnot] Block
Params     ::= Param {"," Param}
Param      ::= Identifier ":" Type
CapAnnot   ::= "!{" CapName {"," CapName} "}"
CapName    ::= Identifier {"." Identifier}

RecordDecl ::= "record" Identifier "{" {FieldDecl ","} "}"
FieldDecl  ::= Identifier ":" Type

EnumDecl   ::= "enum" Identifier "{" {VariantDecl ","} "}"
VariantDecl::= Identifier "{" {FieldDecl ","} "}"

Trailing commas are permitted in Params, FieldDecl lists, and VariantDecl lists.

3.3 Types

Type ::= "Int"
       | "Float"
       | "String"
       | "Bool"
       | "Unit"
       | "Option" "<" Type ">"
       | "Result" "<" Type "," Type ">"
       | "List"   "<" Type ">"
       | "Map"    "<" Type "," Type ">"
       | Identifier                           (* user-declared record or enum *)

There are no anonymous tuples, function types as first-class type expressions, or generic user types in 1.0.

3.4 Statements and expressions

Block      ::= "{" {Stmt Newline} [Expr] "}"
Stmt       ::= LetStmt | AssignStmt | WhileStmt | ForStmt | ExprStmt
LetStmt    ::= "let" ["mut"] Identifier ":" Type "=" Expr
AssignStmt ::= Identifier "=" Expr                       (1.1, §4.5)
WhileStmt  ::= "while" Expr Block                        (1.1, §4.5)
ForStmt    ::= "for" Identifier "in" Expr Block          (1.1, §4.5)
ExprStmt   ::= Expr

Expr       ::= If | Match | Binary | Unary | Call | RecordLit | EnumLit
             | ListLit | MapLit | FieldAccess | Path | Literal | Block
             | "(" Expr ")"

If         ::= "if" Expr Block "else" Block
Match      ::= "match" Expr "{" {MatchArm} "}"
MatchArm   ::= Pattern "=>" Expr [","]

Binary     ::= Expr BinOp Expr
BinOp      ::= "+"|"-"|"*"|"/"|"%"|"=="|"!="|"<"|"<="|">"|">="|"&&"|"||"
Unary      ::= ("-"|"!") Expr

Call       ::= Path "(" [Args] ")"
Args       ::= Expr {"," Expr}

RecordLit  ::= TypeName "{" [Spread ","] {FieldInit ","} "}"
Spread     ::= ".." Expr
FieldInit  ::= Identifier ":" Expr

EnumLit    ::= TypeName "::" Identifier "{" {FieldInit ","} "}"

ListLit    ::= "[" [Expr {"," Expr}] "]"
MapLit     ::= "{" [MapEntry {"," MapEntry}] "}"
MapEntry   ::= Expr ":" Expr

FieldAccess::= Expr "." Identifier
Path       ::= Identifier {"::" Identifier}
TypeName   ::= Identifier
Literal    ::= IntLit | FloatLit | StringLit | BoolLit | UnitLit
             | "Some" "(" Expr ")" | "None"
             | "Ok"   "(" Expr ")" | "Err" "(" Expr ")"

Operator precedence (lowest to highest binding):

1. ||
2. &&
3. == != < <= > >=
4. + -
5. * / %
6. unary - !
7. . :: () (postfix call/field/path)

All binary operators except comparisons are left-associative. Comparisons are non-associative (chaining like a < b < c is rejected).

3.5 Patterns

Pattern    ::= "_"                                 (* wildcard *)
             | Identifier                           (* binding *)
             | Literal
             | "Some" "(" Pattern ")"
             | "None"
             | "Ok"   "(" Pattern ")"
             | "Err"  "(" Pattern ")"
             | TypeName "::" Identifier "{" {FieldPat ","} "}"
FieldPat   ::= Identifier [":" Pattern]

A FieldPat of the form name is shorthand for name: name.

Known gap in the reference implementation: a negative integer literal (-1) is not accepted as a pattern (parse error “expected pattern, found Minus”). Use a guard-free alternative such as an if on the value until this is supported.

4. Type system

Types: Int, Float, String, Bool, Unit, Option<T>, Result<T, E>, List<T>, Map<K, V>, user-declared records, user-declared enums.

There are no user-definable generic types, no traits, no subtyping. Every type is concrete at definition time. Type equivalence is structural for built-in generics (List<Int> ≡ List<Int>) and nominal for records and enums.

4.1 Built-in type rules (judgments)

We use Γ ⊢ e : T for “in environment Γ, expression e has type T”. Γ maps identifiers to types and tracks the in-scope capability set Φ (see §6).

Literals:

─────────────────────                ─────────────────────
Γ ⊢ IntLit n   : Int                Γ ⊢ FloatLit f : Float

─────────────────────                ─────────────────────
Γ ⊢ StringLit s: String             Γ ⊢ BoolLit b  : Bool

─────────────────────
Γ ⊢ ()         : Unit

Variables:

   x : T  ∈  Γ
─────────────────────
Γ ⊢ x : T

Let bindings:

Γ ⊢ e : T          T = AnnotatedType
────────────────────────────────────────
Γ, x : T  ⊢   following statements

A let annotation MUST exactly match the inferred type of the right-hand side (no coercion, no widening). 1.0 has no implicit conversions between Int and Float.

If-else:

Γ ⊢ cond : Bool       Γ ⊢ b1 : T       Γ ⊢ b2 : T
────────────────────────────────────────────────────
Γ ⊢ if cond { b1 } else { b2 } : T

Both branches MUST exist and have the same type. There is no single-branch if as an expression.

Arithmetic:

Γ ⊢ a : T   Γ ⊢ b : T    T ∈ {Int, Float}
─────────────────────────────────────────
Γ ⊢ a + b : T            (same for - * / %)

% is defined for Int only.

Comparison:

Γ ⊢ a : T   Γ ⊢ b : T    T ∈ {Int, Float, String, Bool}
─────────────────────────────────────────────────────────
Γ ⊢ a == b : Bool        (same for !=)

Γ ⊢ a : T   Γ ⊢ b : T    T ∈ {Int, Float, String}
─────────────────────────────────────────────────
Γ ⊢ a <  b : Bool        (same for <= > >=)

Logical:

Γ ⊢ a : Bool   Γ ⊢ b : Bool                 Γ ⊢ a : Bool
────────────────────────────                ──────────────
Γ ⊢ a && b : Bool   (same for ||)           Γ ⊢ !a : Bool

String concatenation: The + operator on two String values yields a String. Mixed-type concatenation is rejected.

4.2 Compound types

Option / Result / List / Map literals:

Γ ⊢ e : T                                Γ ⊢ e : T
──────────────────────                   ──────────────────────
Γ ⊢ Some(e) : Option<T>                  Γ ⊢ Ok(e)  : Result<T, _>

                                         Γ ⊢ e : E
─────────────────                        ──────────────────────
Γ ⊢ None : Option<_>                     Γ ⊢ Err(e) : Result<_, E>

Γ ⊢ e1 : T  ...  Γ ⊢ en : T               Γ ⊢ k_i : K   Γ ⊢ v_i : V (for all i)
─────────────────────────────             ─────────────────────────────────
Γ ⊢ [e1, ..., en] : List<T>                Γ ⊢ {k_1: v_1, ...} : Map<K, V>

The element type of None, Ok, Err, [], and {} is determined by the surrounding context (let annotation or function return type).

Records:

record R { f1: T1, ..., fn: Tn }   ∈   Γ
Γ ⊢ e1 : T1   ...   Γ ⊢ en : Tn
─────────────────────────────────────────
Γ ⊢ R { f1: e1, ..., fn: en } : R

Record spread:

Γ ⊢ base : R     R = record { ..., fk: Tk, ... }
Γ ⊢ ek_new : Tk   (for each overridden field)
────────────────────────────────────────────────
Γ ⊢ R { ..base, fk: ek_new, ... } : R

Spread rules:

  1. The base expression MUST have type R, the same record type as the literal.
  2. Overrides MAY override any subset of fields, including none.
  3. A field MUST NOT be specified more than once.
  4. Spread MUST appear first inside the braces; non-spread fields override.
  5. The result is a fresh record value with the original base unchanged (records are immutable, §7.2).

Enums:

enum E { V { f1: T1, ..., fn: Tn }, ... }   ∈   Γ
Γ ⊢ e1 : T1   ...   Γ ⊢ en : Tn
─────────────────────────────────────────────
Γ ⊢ E::V { f1: e1, ..., fn: en } : E

4.3 Functions and calls

fn f(p1: T1, ..., pn: Tn) -> R [!{φ_f}]   ∈   Γ
Γ ⊢ a1 : T1   ...   Γ ⊢ an : Tn
φ_f ⊆ current capability set Φ
─────────────────────────────────────────────────
Γ ⊢ f(a1, ..., an) : R

Function bodies type-check with the parameters in scope. The last expression of the block determines the return type and MUST equal R.

4.4 Pattern matching

Match expressions select an arm whose pattern matches the scrutinee. All arms MUST produce the same type.

Γ ⊢ s : T     for each arm i:  Pi covers some subset of T
                                Γ, bindings(Pi) ⊢ ei : U
arms are exhaustive over T
────────────────────────────────────────────────────────
Γ ⊢ match s { P1 => e1, ..., Pn => en } : U

4.4.1 Exhaustiveness

A match is exhaustive when every value of the scrutinee type is matched by at least one arm. The compiler MUST reject non-exhaustive matches.

Exhaustiveness rules per scrutinee type:

Scrutinee typeExhaustive iff
Boolboth true and false arms, OR a binding/wildcard arm
Option<T>both Some(_) and None arms, OR a binding/wildcard arm
Result<T, E>both Ok(_) and Err(_) arms, OR a binding/wildcard arm
enum Eevery variant of E, OR a binding/wildcard arm
Int, Float, Stringa binding or wildcard arm is REQUIRED
record Ra binding or wildcard arm is REQUIRED (records are not destructured exhaustively in 1.0)
Unittrivially exhaustive (the only inhabitant is ())

A wildcard _ or a bare-identifier binding always makes a match exhaustive.

4.4.2 Reachability

The compiler SHOULD warn on unreachable arms (an arm whose pattern is fully covered by an earlier arm). 1.0 does not require this warning to be a hard error.

4.5 Rebinding and loops (added in 1.1)

Mutable bindings and assignment. let mut x: T = e introduces a binding that MAY later be rebound with x = e2. Assignment changes which value the name x refers to for the rest of its scope. It never changes a value: records, lists and maps are still never updated in place (§7.2), so any other name holding the old value still sees the old value.

x : T (declared mut) ∈ Γ      Γ ⊢ e : T
──────────────────────────────────────────
Γ ⊢ x = e   ok
  • The assigned value MUST have the binding’s type. The reference implementation does not check this yet; a program that assigns a value of another type has unspecified behaviour.
  • Rebinding a binding declared without mut, a function parameter or a for loop variable is not permitted by this specification. The reference implementation still accepts it, because 1.0 programs (including two standard libraries) relied on it, and §1.2 forbids tightening type rules within 1.x. Tooling reports it as warning E010 with an automatic fix (boruna lang repair adds mut); it becomes a compile error in language version 2.0.
  • An assignment inside a nested block (if, match arm, loop body) rebinds the binding from the enclosing scope. A let inside a block introduces a new binding that ends with the block.

While loops. while cond { body } evaluates cond; while it is true it runs body and evaluates cond again. cond MUST have type Bool (the reference implementation does not check this yet). A while statement has no value.

For loops. for v in e { body } evaluates e once. e MUST have type List<T>; iterating anything else is a runtime error in the reference implementation. body runs once per element in list order, with v : T bound to the element. v and any let inside body are scoped to the body. A for statement has no value.

Termination. Loops do not change the determinism model (§7): the same inputs run the same iterations. A loop that never ends is stopped by the VM’s step limit (--step-limit), which reports a runtime error rather than hanging.

5. Pattern binding

Pattern matching introduces bindings into the arm’s scope:

  • _ introduces no binding.
  • An Identifier pattern binds the matched value under that name.
  • Some(P), Ok(P), Err(P) recursively bind from P.
  • E::V { f1: P1, f2, ... } binds each Pi. The shorthand f2 binds the value of field f2 to the identifier f2.

Bindings are only in scope within the arm’s right-hand side.

5a. Standard built-in functions

The following functions are provided by the runtime in every compilation unit. They require no import. Their names begin with __builtin_ to prevent shadowing by user-defined identifiers.

All built-ins are pure (no capability annotation). Their semantics are defined below.

String operations

NameSignatureSemantics
__builtin_int_to_string(Int) -> StringReturns the decimal string representation of the argument.
__builtin_float_to_string(Float) -> StringReturns a string representation of the float argument.
__builtin_string_len(String) -> IntReturns the number of bytes in the UTF-8 encoding of the string.
__builtin_string_chars(String) -> List<String>Returns a list of single-character strings, one per Unicode scalar value.
__builtin_string_contains(String, String) -> BoolReturns true iff the first argument contains the second as a substring.
__builtin_string_starts_with(String, String) -> BoolReturns true iff the first argument begins with the prefix given by the second.
__builtin_string_ends_with(String, String) -> BoolReturns true iff the first argument ends with the suffix given by the second.
__builtin_string_to_upper(String) -> StringReturns a copy of the string with all ASCII alphabetic characters uppercased.
__builtin_string_to_lower(String) -> StringReturns a copy of the string with all ASCII alphabetic characters lowercased.
__builtin_string_trim(String) -> StringReturns a copy of the string with leading and trailing ASCII whitespace removed.
__builtin_string_join(List<String>, String) -> StringReturns the elements of the list concatenated, with the second argument inserted between each pair of adjacent elements.

List operations

NameSignatureSemantics
__builtin_list_len(List<T>) -> IntReturns the number of elements in the list.
__builtin_list_is_empty(List<T>) -> BoolReturns true iff the list has zero elements.
__builtin_list_head(List<T>) -> Option<T>Returns Some(first) if the list is non-empty, otherwise None.
__builtin_list_tail(List<T>) -> List<T>Returns a new list containing all elements after the first. Returns an empty list if the argument is empty.
__builtin_list_append(List<T>, T) -> List<T>Returns a new list equal to the original with the second argument appended at the end.
__builtin_list_concat(List<T>, List<T>) -> List<T>Returns a new list that is the concatenation of the two arguments, in order.
__builtin_list_reverse(List<T>) -> List<T>Returns a new list containing the same elements in reversed order.

All list built-ins are non-mutating; the original list is unchanged. This is consistent with the immutability requirement in §7.2.

6. Capability semantics

6.1 Annotation form

A function declaration MAY include a capability annotation after the return type:

fn f(...) -> R !{cap_1, cap_2, ..., cap_n}

Each cap_i is a dotted name like net.fetch. The annotation is the declared capability set of the function, written φ_f.

A function without an annotation has φ_f = ∅.

6.2 The capability namespace (1.0 frozen set)

The following capabilities are defined in 1.0. Implementations MUST recognize each and MUST NOT silently rename them:

NameNumeric IDDescription
net.fetch0Network requests
fs.read1File system read
fs.write2File system write
db.query3Database queries
ui.render4UI rendering
time.now5Current time
random6Random number generation
llm.call7LLM provider call
actor.spawn8Spawn an actor
actor.send9Send an actor message
step.input10Read a workflow step’s resolved input value

Implementations MAY add new capabilities at IDs ≥ 11 in future 1.y releases; the IDs and names above are frozen.

6.3 Propagation rule

For any call g(...) made within the body of f:

φ_g  ⊆  φ_f

The required capability set of every callee MUST be a subset of the caller’s declared set. Equivalently: a function may only invoke effects it has declared.

This is checked statically at type-check time. The reference implementation rejects programs that violate this with a CapabilityError.

6.4 Composition

Capability sets compose as set union along call edges:

φ_caller_required = ⋃ { φ_g : g is reachable from caller }

For a top-level entry point (e.g. main), the union of φ_g over all reachable g is the effective capability set of the program.

6.5 Runtime gating

At runtime, the host VM holds an active policy Π that is a subset of the capability namespace. When bytecode invokes a capability c:

  1. The VM checks c ∈ Π. If not, the call fails with a capability-denied error.
  2. The VM dispatches to the registered handler for c.
  3. The result (or error) is recorded in the event log for replay (§7.3).

Capability annotations are static facts about the program; the policy Π is a runtime fact about the deployment. The two MUST agree at execution time.

6.6 No ambient capabilities

There are no implicit, ambient, or hidden capabilities. Every effect MUST be declared in φ_f and MUST be present in Π. Functions without annotations are pure with respect to the capability set; they MUST NOT directly invoke any capability call.

7. Determinism

7.1 Definition

For any program P and any input I, executing P(I) MUST produce the same observable result on every run, given the same recorded capability outcomes (§7.3).

7.2 Immutability

All values are immutable. There is no mutable cell and no in-place update. Record spread (§4.2) constructs a new value. Since 1.1 a let mut binding can be rebound to a different value (§4.5); that changes what the name refers to, never the value itself, and is local to one function call, so it does not affect determinism.

7.3 Replay model

Side effects are recorded by the host VM at the boundary defined in §6.5. A recorded run produces an event log of capability invocations and outcomes. Replaying the program against the same event log MUST produce the same output, byte-for-byte (subject to the implementation’s serialization choices for Map ordering — see §7.4).

7.4 Map iteration order

Map<K, V> iteration order is insertion order in 1.0. The reference implementation uses a BTreeMap, which yields key-sorted iteration. A 1.0-conformant implementation MUST choose one and document it; both choices are observationally indistinguishable for inputs that do not depend on iteration order, and 1.x will not break programs that rely on either choice.

Implementations SHOULD document which order they provide. The reference implementation is key-sorted.

7.5 Forbidden in pure code

Pure (non-capability-annotated) code MUST NOT:

  • Read the system clock
  • Generate randomness
  • Read from / write to the file system or network
  • Depend on memory addresses, hash randomization seeds, or thread scheduling

If any of these are needed, they MUST go through a capability call (§6).

8. Standalone programs and entry points

8.1 Executable programs

A .ax source file intended as a standalone program MUST declare:

fn main() -> Int { ... }

The integer return value is the program’s exit code. 0 is success.

8.2 Framework apps

A framework app (Boruna’s Elm-architecture protocol) MUST declare three top-level functions:

fn init() -> State
fn update(state: State, msg: Msg) -> UpdateResult
fn view(state: State) -> UINode

with user-declared types State, Msg, Effect, UpdateResult, UINode, and PolicySet. The framework protocol layered on .ax is documented separately in docs/concepts/ and is not part of the language spec proper.

9. Source examples (informative)

The following are well-formed 1.0 programs.

Hello, sum:

fn add(a: Int, b: Int) -> Int {
    a + b
}

fn main() -> Int {
    add(2, 40)
}

Record spread:

record Point {
    x: Int,
    y: Int,
}

fn shift_y(p: Point, dy: Int) -> Point {
    Point { ..p, y: p.y + dy }
}

fn main() -> Int {
    let p: Point = Point { x: 3, y: 4 }
    let q: Point = shift_y(p, 10)
    q.y
}

Capability-annotated function (compile-only — runtime denied without net.fetch policy):

fn fetch(url: String) -> String !{net.fetch} {
    "stub"
}

fn main() -> Int {
    let body: String = fetch("https://example.com")
    0
}

Match exhaustiveness over an enum:

enum Shape {
    Circle    { radius: Float },
    Rectangle { width: Float, height: Float },
}

fn label(s: Shape) -> String {
    match s {
        Shape::Circle    { radius } => "circle"
        Shape::Rectangle { width, height } => "rectangle"
    }
}

fn main() -> Int {
    0
}

10. Errors and diagnostics (informative)

The reference implementation surfaces errors at three layers — lexer, parser, and type checker — via the CompileError enum. Conformant implementations are not required to use the same error categories, but SHOULD distinguish:

  • Lexical errors (unterminated string, invalid escape, illegal character)
  • Parse errors (unexpected token, missing delimiter)
  • Type errors (mismatched types, undefined identifier, non-exhaustive match, capability propagation violation, duplicate field, missing field, invalid spread)

11. Bytecode mapping (informative)

.ax source compiles to Boruna bytecode (.axbc). The bytecode opcode set, capability IDs, and binary format are specified in docs/bytecode-spec.md. The capability ID table in §6.2 is frozen jointly with the bytecode spec for 1.x.

12. Cross-references

13. Change log for this specification

  • 1.0 (2026-04-28) — Initial freeze. Sprint W1-B. Captures the language as shipped in Boruna v0.5.0.
  • 1.1 (2026-10-03) — Additive (§1.2). Specifies what the implementation has accepted since Boruna v2.0: let mut, assignment, while and for (§4.5); mut, while, for and in become keywords. Rebinding a binding that is not mut stays accepted but is warning E010.

Bytecode


spec_id: bytecode bytecode_version: “1.0” status: stable last_revised: 2026-04-28 sprint: W9-A audience: VM implementers, compiler authors, replay/evidence tooling

Boruna Bytecode Specification — Version 1.0

This document is the formal specification of the Boruna bytecode format (.boruna_bytecode / .axbc). It is the authoritative reference for any independent implementation of a Boruna VM, replay engine, or evidence verifier.

The narrative, example-driven companion is docs/bytecode-spec.md. When the two disagree, this spec wins.

The reference implementation lives in crates/llmbc/ (Op, Value, Capability, Module) and crates/llmvm/ (the bytecode interpreter) of the Boruna repository. Implementations against this spec MAY use a different code organization but MUST be observationally equivalent to the rules below.

1. Conformance and versioning

1.1 Version identifier

The current bytecode version is 1.1. Implementations MUST expose this value programmatically.

In the reference implementation:

#![allow(unused)]
fn main() {
// crates/llmbc/src/lib.rs
pub const BYTECODE_VERSION: &str = "1.1";
}

The spec version 1.1 is the public, semver-like format identifier. The on-disk module header carries an internal version byte (currently 1, see §3.1) which is incremented for any wire-format change inside the 1.x line; clarifying spec edits do not bump it.

1.1 (additive minor bump, this session) adds two opcodes: Op::Debug at byte tag 0xA7 and Op::DebugMsg at byte tag 0xA8 (see §4.5). Per §1.2(6) these are additive only; a 1.0 reader presented with either MUST reject with a typed unknown-opcode error.

The BYTECODE_VERSION string is a <major>.<minor> decimal number. A bytecode module emitted against 1.x MUST load and execute against any 1.y VM where y >= x.

1.2 Backwards-compatibility commitment for 1.x

Within the 1.x line:

  1. Trace-and-output stability. Any 1.0 bytecode module loaded by any 1.x VM MUST produce the same trace (event log) and the same output value, given the same recorded capability outcomes.
  2. Opcode discriminants frozen. The byte tags assigned in §4.2 are locked. Reusing a tag for a different opcode is a 2.0 break.
  3. Capability IDs frozen. The (id, name) pairs in §6.1 are locked for 1.x. Reordering or renaming is a 2.0 break.
  4. Value variants frozen. The set and shape of Value variants in §5.1 are locked. New variants are a 2.0 break (existing modules MAY contain them only if every reachable VM supports them — i.e. additive only at a 1.y minor bump with reader gating).
  5. Module header layout frozen. The magic bytes, version field width, and length-prefixed payload structure (§3.1) are locked.
  6. Additive opcodes are a minor bump. A 1.y y > 0 MAY introduce new opcodes at currently-unassigned discriminants. A 1.0 reader presented with such a module MUST reject the module with a typed unknown-opcode error rather than guess. (See §10.)

Breaking changes (renames, removals, discriminant reuse, header reshuffles) are deferred to 2.0 or later.

1.3 Conformance levels

A 1.0 VM implementation MUST:

  • Accept any module whose header magic and bytecode_version match §3.1 and whose internal payload deserializes against §3.2.
  • Execute every opcode in §4 with the stack/operand effects and behavior described.
  • Enforce capability gating (§6) against the active host policy at every CapCall.
  • Preserve determinism (§7) for all pure code (code paths not invoking CapCall, SpawnActor, SendMsg, or ReceiveMsg).
  • Reject unknown opcodes, unknown capability IDs, and bytecode_version >= 2 modules with a typed error (§10).

A 1.0 VM implementation MAY:

  • Provide additional debugging, tracing, or profiling instrumentation, as long as it does not change observable execution behavior.
  • Implement opcodes via JIT or AOT translation, as long as the stack effect and outputs match.

2. Document conventions

  • “Reader” means any consumer of a bytecode module (VM, evidence verifier, disassembler).
  • “Writer” means any producer of a bytecode module (the compiler, codegen tools).
  • “MUST” / “MUST NOT” / “SHOULD” / “MAY” follow RFC 2119 usage.
  • All multi-byte integers are little-endian unless explicitly stated.
  • “Stack effect” describes operand stack changes during execution (not constant-pool or local-frame changes).

3. Module format

A bytecode module is the unit of distribution. Each module is a self-contained graph of functions, types, constants, and globals.

3.1 On-disk header

Modules are persisted in a length-prefixed binary container:

OffsetSize (bytes)FieldNotes
04magicExactly the four bytes 0x4C 0x4C 0x4D 0x42 ("LLMB").
42versionLittle-endian u16. 0x0001 in 1.0.
64payload_lengthLittle-endian u32. Byte length of the payload that follows.
10payload_lengthpayloadUTF-8 JSON encoding of the Module struct (§3.2).

Readers MUST:

  • Reject any input shorter than 10 bytes with a typed error.
  • Reject any input whose first four bytes do not equal the magic.
  • Reject any input whose version field is not 0x0001 with an UnsupportedVersion error. This is the §1 application: reject across-major bytecode_version, never silently accept.
  • Reject any input where data.len() < 10 + payload_length with a truncation error.

3.2 Payload encoding

The payload is canonical JSON of the Module shape. Field order in the on-disk representation matches Rust struct field order; readers MUST tolerate any order (JSON object semantics).

{
  "name": "<module name>",
  "version": 1,
  "constants": [<Value>, ...],
  "globals": ["<global name>", ...],
  "types":   [<TypeDef>, ...],
  "functions": [<Function>, ...],
  "entry": <function index, u32>
}

The reference encoder uses serde_json::to_vec (no whitespace control; readers MUST NOT depend on a specific JSON whitespace style). The reference decoder uses serde_json::from_slice.

Encoding choice (informative). 1.0 specifies a JSON payload wrapped in the binary header. This choice trades raw bytes-per-module for human-debuggability and stable cross-language deserialization. A future major (2.0) MAY swap to a denser binary encoding (CBOR, postcard, custom); until then, the JSON contract is part of the surface.

3.3 Sections

The payload contains five logical sections:

  • name — UTF-8 module identifier. Used for diagnostics; not security-load-bearing.
  • version — Internal u16 version of the wire format; currently 1. This is distinct from the spec bytecode_version (§1.1).
  • constants — The constant pool. Indexed by PushConst(idx). Each entry is a Value (§5).
  • globals — Names of module-level globals. Indexed by LoadGlobal(idx) / StoreGlobal(idx).
  • types — Type definitions for records and enums. See §3.4. Indexed by the type_id operand of MakeRecord, MakeEnum, and the type_id field of Value::Record / Value::Enum.
  • functions — The function table. Each entry is a Function. Indexed by Call(fn_idx, _), SpawnActor(fn_idx), and Value::FnRef(idx).
  • entry — The index into functions of the program entry point (main).

There is no separate “capability table” in the wire format; declared capabilities live on each Function (§3.4) as a Vec<Capability>. The frozen capability namespace is in §6.

3.4 Function and type shapes

Function {
  name: String,
  arity: u8,
  locals: u16,
  code: Vec<Op>,
  capabilities: Vec<Capability>,
  match_tables: Vec<Vec<MatchArm>>,
}

MatchArm {
  tag: i32,        // variant index or literal tag; -1 = wildcard
  target: u32,     // jump offset within the function's code
}

TypeDef {
  name: String,
  kind: TypeKind,
}

TypeKind ::=
    Record { fields: Vec<(String, String)> }      // (field name, type name)
  | Enum   { variants: Vec<(String, Option<String>)> }  // (variant name, optional payload type name)

Field invariants:

  • arity is the number of arguments the function expects on the stack at entry. The VM pops them into locals 0..arity.
  • locals is the total local slot count (including parameters).
  • capabilities lists the static capability set declared by this function (§6).
  • match_tables are referenced by Match(table_idx) opcodes; each entry is the arm list for one match.

4. Opcodes

The VM is stack-based. Each function body is a flat sequence of Ops. The operand stack and the local frame are private to each function activation.

4.1 Stack effect notation

(a, b → c) reads as: pops a then b (so b is on top), pushes c. Order matters; the stack-top operand is the rightmost in the “before” list.

(→ x) pushes x onto the stack; (x →) pops x from the stack.

4.2 Opcode table

Every opcode below is part of the frozen 1.0 set. The byte tag column gives the discriminant used by Op::to_byte_tag() in the reference implementation; tags are locked for 1.x.

OpcodeByte tagStack effectBehavior
PushConst(idx)0x01(→ v)Push constants[idx]. Deterministic.
LoadLocal(idx)0x02(→ v)Push locals[idx] from the current frame. Deterministic.
StoreLocal(idx)0x03(v →)Pop and store into locals[idx]. Deterministic.
LoadGlobal(idx)0x04(→ v)Push globals[idx] from the module-level globals. Deterministic.
StoreGlobal(idx)0x05(v →)Pop and store into globals[idx]. Deterministic.
Call(fn_idx, arity)0x06(a₁..aₙ → r)Pop arity args (left-to-right pushed → rightmost on top), call functions[fn_idx], push return value.
Ret0x07(v →)Return from current function. Pops return value from the callee stack and pushes it on the caller stack.
Jmp(off)0x08(—)Unconditional jump to instruction offset off within the current function.
JmpIf(off)0x09(cond →)Pop; if cond.is_truthy() (§5.2), jump to off. Otherwise fall through.
JmpIfNot(off)0x0A(cond →)Pop; if NOT cond.is_truthy(), jump to off. Otherwise fall through.
Match(table_idx)0x0B(v → …)Pop scrutinee; consult match_tables[table_idx]. First arm whose tag matches: jump to arm.target. Fallthrough on wildcard (tag == -1).
MakeRecord(type_id, n)0x0C(f₁..fₙ → r)Pop n field values, push Value::Record { type_id, fields } with fields in pop-order reversed.
MakeEnum(type_id, var)0x0D(payload → e)Pop one payload value (or Unit), push Value::Enum { type_id, variant: var, payload }.
GetField(idx)0x0E(r → v)Pop a Record, push the field at index idx.
SpawnActor(fn_idx)0x0F(→ id)Spawn an actor running functions[fn_idx]; push Value::ActorId(_). Capability-gated by actor.spawn.
SendMsg0x10(target_id, msg →)Pop message and target id; deliver to actor mailbox. Capability-gated by actor.send.
ReceiveMsg0x11(→ msg)Block current actor until a message arrives; push it.
Assert(err_idx)0x12(v →)Pop; if !v.is_truthy(), abort with constants[err_idx] as the error. Otherwise no-op.
CapCall(cap_id, n)0x13(a₁..aₙ → r)Pop n args, invoke capability cap_id (§6) via the host gateway, push the result. Capability-gated.
Add0x20(a, b → a+b)Numeric add. Both operands MUST be Int+Int or Float+Float.
Sub0x21(a, b → a−b)Numeric subtract.
Mul0x22(a, b → a·b)Numeric multiply.
Div0x23(a, b → a/b)Numeric divide. Int / 0 traps; Float / 0.0 follows IEEE 754 (±inf / NaN).
Mod0x24(a, b → a%b)Integer remainder. Defined for Int only; Mod 0 traps.
Neg0x25(a → −a)Numeric negation.
Eq0x30(a, b → Bool)Structural equality (§5.3).
Neq0x31(a, b → Bool)Negated Eq.
Lt0x32(a, b → Bool)Total ordering for Int, Float, String (§5.3). Other types: VM error.
Lte0x33(a, b → Bool)Same domain as Lt.
Gt0x34(a, b → Bool)Same domain as Lt.
Gte0x35(a, b → Bool)Same domain as Lt.
Not0x40(a → Bool)Boolean negation; argument MUST be Bool.
And0x41(a, b → Bool)Boolean AND. Eager (both operands evaluated by codegen before the opcode).
Or0x42(a, b → Bool)Boolean OR. Eager.
Concat0x50(s1, s2 → s)UTF-8 string concatenation; both operands MUST be String.
Pop0x60(v →)Discard top of stack.
Dup0x61(v → v, v)Duplicate top of stack. Reference-counted clone for owned data.
EmitUi0x70(tree → tree)Emit a UI descriptor (top of stack) to the host UI sink. Stack effect is conventionally the identity (the tree is left on the stack for chaining); see §4.3. Capability-gated by ui.render.
MakeList(n)0x80(v₁..vₙ → list)Pop n values (rightmost on top), push a List containing them in pop-reversed order.
ListLen0x81(list → Int)Pop a List, push its length.
ListGet0x82(list, idx → v)Pop list and index. If 0 <= idx < len, push the element. Out-of-bounds traps.
ListPush0x83(list, v → list’)Pop list and value; push a NEW list (immutable; original unchanged) with the value appended.
ParseInt0x84(String → Int)Pop a String, parse as decimal i64; push 0 on parse failure. Determinism preserved.
StrContains0x85(haystack, needle → Bool)Pop two Strings; push Bool.
StrStartsWith0x86(string, prefix → Bool)Pop two Strings; push Bool.
TryParseInt0x87(String → Result<Int, String>)Pop a String; push Ok(Int) on success, Err(String) on failure.
Nop0xFE(—)No effect. Useful for jump targets.
Halt0xFF(—)Halt VM execution. The top of stack is the program result.

4.3 Per-opcode semantic notes

  • EmitUi’s observable effect is the rendered tree. The VM treats it as identity on the stack; the host UI sink is invoked with the tree as a side effect at the ui.render capability boundary.
  • Call records a return pointer in the call stack; Ret returns to it. The depth of the call stack is host-bounded.
  • Match arms are tried in order; the first match wins. A tag == -1 arm is unconditional and acts as the wildcard.
  • SpawnActor, SendMsg, and ReceiveMsg interact with the host scheduler. The scheduler ordering MUST be deterministic (round-robin sorted by (target_id, sender_id); see docs/concepts/determinism.md).

4.4 Reserved opcode space

The byte tag space contains gaps (e.g. 0x14–0x1F, 0x26–0x2F, 0x36–0x3F, 0x43–0x4F, 0x51–0x5F, 0x62–0x6F, 0x71–0x7F, 0x88–0xFD). These are reserved for additive opcodes in 1.y minor bumps. A 1.0 reader MUST reject any encountered tag not listed in §4.2 with an unknown-opcode error.

4.5 1.1 additions (debug print-and-passthrough)

The following opcodes were introduced in bytecode version 1.1 per §1.2(6). They consume tag space previously reserved under §4.4. A 1.0 reader MUST reject any module containing either with an unknown-opcode error.

OpcodeByte tagStack effectBehavior
Debug0xA7(v → v)Pop one value; write its Value::Display form followed by a newline to host stderr; push the value back unchanged. No capability gate; no audit-log event; never feeds replay verification. Operational-only (§8.2).
DebugMsg0xA8(msg, v → v)Pop value, then pop message; write <msg> <value>\n to host stderr (the message is Display-formatted if not already a String); push value back unchanged. Same operational-only semantics as Debug.

Both opcodes are stack-effect identity for the value flow — they exist only for human-driven debugging of .ax source. Compiler surface: __builtin_debug(v) and __builtin_debug_msg(msg, v). See docs/architecture-q-debug.md for the design rationale.

Determinism note. A workflow that calls Debug or DebugMsg is still deterministic in its replay-verified state (§8.1) because the value flow is identity. The stderr output is operational-only (§8.2). Evidence-bundle hash chains are unaffected.

5. Value model

5.1 Value variants

The frozen Value discriminant set in 1.0:

VariantPayloadType tag (type_name())Notes
Unit(none)"Unit"The single inhabitant of the Unit type.
Bool(b)bool"Bool"
Int(n)i64"Int"64-bit signed two’s complement.
Float(f)f64"Float"IEEE 754 double. See §7 for determinism caveats.
String(s)UTF-8 String"String"
None(none)"None"Option::None.
Some(v)Box<Value>"Some"Option::Some.
Ok(v)Box<Value>"Ok"Result::Ok.
Err(v)Box<Value>"Err"Result::Err.
Record{ type_id: u32, fields: Vec<Value> }"Record"Field order matches the corresponding TypeDef::Record::fields order.
Enum{ type_id: u32, variant: u8, payload: Box<Value> }"Enum"Variant index matches TypeDef::Enum::variants order.
List(items)Vec<Value>"List"Insertion-ordered, immutable; ListPush returns a new list.
Map(entries)BTreeMap<String, Value>"Map"Always key-sorted in 1.0. See §7.3.
ActorId(id)u64"ActorId"Opaque actor handle. Compared by id.
FnRef(idx)u32"FnRef"Function-table index. Used for higher-order references.

This set is frozen for 1.x. Adding a variant is a 2.0 break.

5.2 Truthiness

The is_truthy(v) predicate, used by JmpIf, JmpIfNot, and Assert:

VariantTruthy iff
Unitalways false
Bool(b)b == true
Int(n)n != 0
Float(f)f != 0.0 (note: NaN != 0.0 is true under IEEE)
String(s)!s.is_empty()
Nonealways false
Some(_)always true
Ok(_)always true
Err(_)always false
Recordalways true
Enumalways true
List(l)!l.is_empty()
Map(m)!m.is_empty()
ActorId(_)always true
FnRef(_)always true

5.3 Equality and ordering

  • Equality (Eq / Neq) is structural: same variant + same payload (recursively, with byte-equal String, IEEE-bit-equal Float, structural Vec and BTreeMap). NaN != NaN (IEEE); writers SHOULD avoid relying on Float equality.
  • Ordering (Lt, Lte, Gt, Gte) is defined for Int (signed numeric), Float (IEEE; NaN comparisons are always false), and String (lexicographic UTF-8 byte order). Ordering on other types is a runtime error.
  • The discriminant order in §5.1 is not an ordering relation; cross-variant comparisons (e.g. Int < String) trap.

5.4 In-memory representation

The reference VM holds Values as boxed/refcounted handles (Rust Box/Vec/String/BTreeMap). Implementations MAY use any representation that preserves the equality, ordering, and truthiness contracts above.

5.5 Serialization to evidence-bundle JSON

Evidence bundles serialize Values as canonical JSON via the same serde::Serialize derive used for the bytecode payload (§3.2). The JSON surface for each variant mirrors the Rust enum’s tagged form (e.g. Some(v) → {"Some": <v>}; Ok(v) → {"Ok": <v>}; Record → {"Record": {"type_id": ..., "fields": [...]}}). Because Map is a BTreeMap, the JSON object key order in the bundle is always sorted by key — this is a load-bearing determinism contract (§7.3) for workflow_hash and replay verification.

6. Capability table

6.1 Frozen 1.0 capability namespace

Implementations MUST recognize each capability below by both its numeric ID and its wire string. The (id, name) pairs are locked for 1.x. Reordering is a 2.0 break.

CapabilityNumeric IDWire nameReplay-verified?Description
NetFetch0net.fetchyesHTTP request; result recorded in event log.
FsRead1fs.readyesRead file contents.
FsWrite2fs.writeyesWrite file contents.
DbQuery3db.queryyesDatabase query.
UiRender4ui.renderoperationalEmit a UI tree to the host sink.
TimeNow5time.nowyesRead wall clock; replay reproduces the recorded value.
Random6randomyesGenerate randomness; replay reproduces recorded bytes.
LlmCall7llm.callyesLLM provider invocation.
ActorSpawn8actor.spawnyesSpawn an actor.
ActorSend9actor.sendyesSend a message to an actor.
StepInput10step.inputyesRead an upstream workflow step’s resolved input.

The 1.0 capability set has exactly 11 entries. Implementations MAY add new capabilities at IDs ≥ 11 in future 1.y releases; previously-assigned IDs and names MUST NOT change.

The Capability::ALL constant in the reference implementation enumerates these in canonical sorted-by-name order (locked by tests::test_capability_all_is_sorted_by_name). This sort order feeds compute_capability_set_hash (see docs/reference/capability-identity.md).

6.2 Capability semantics

  • Static side. A Function’s capabilities: Vec<Capability> declares the set of capabilities its body MAY directly invoke. The compiler enforces propagation (every callee’s set is a subset of the caller’s; see docs/spec/ax-language-1.0.md §6.3).
  • Runtime side. At each CapCall, the host gateway checks the capability id against the active policy Π. If absent, the call fails with a capability-denied error. If present, the gateway dispatches to the registered handler.
  • Recording. Replay-verified capabilities have their (request, result) pair appended to the event log. Operational capabilities (currently only ui.render) emit observable side effects but are NOT replayed for output verification.

6.3 Capability contract version

Each capability also carries a contract version: &'static str (currently "1" for all 11). The version pins the call/return shape and side-effect semantics; it is bumped only on contract-breaking changes (not on every binary release). The combined (name, version) set is hashed by compute_capability_set_hash and exposed via the CapabilitySetReport wire structure. See docs/reference/capability-identity.md for the byte-exact hash algorithm.

7. Determinism contract

The bytecode VM promises that two runs of the same module against the same recorded capability outcomes produce byte-identical observable output, traces, and evidence-bundle hashes.

7.1 Pure code

Code paths that do not invoke CapCall, SpawnActor, SendMsg, ReceiveMsg, or EmitUi are deterministic by construction:

  • All Values are immutable.
  • All control flow is data-driven from the operand stack and locals.
  • No opcode reads the system clock, randomness, or environment.
  • No opcode depends on memory addresses, hash randomization, or thread scheduling.

7.2 Floating-point

Float arithmetic uses IEEE 754 double precision. The bytecode VM does NOT reorder operations, fuse multiply-adds, or use higher-precision intermediates. The same module on different hardware MUST produce identical bit patterns for Add, Sub, Mul, Div, and Neg over Float operands.

NaN payloads are NOT canonicalized; bit patterns from upstream computations propagate unchanged. Writers SHOULD treat NaN as a non-deterministic source and avoid hashing or logging it.

7.3 Map iteration order

Value::Map is a BTreeMap<String, Value>. Iteration order is always key-sorted, ascending byte-lexicographic on UTF-8 keys. This is locked for 1.x:

  • Evidence-bundle JSON serialization uses this order.
  • workflow_hash and bundle hashes depend on this order.
  • Any 1.x VM that diverges from key-sorted Map iteration is non-conformant.

(Compare: .ax source language §7.4 leaves implementations a choice between insertion order and key-sorted; the bytecode level is stricter — only key-sorted is conformant.)

7.4 Actor scheduling

The actor scheduler is deterministic: round-robin over actors sorted by (target_id, sender_id). The scheduler tick, message-send, and message-receive events are all logged in the EventLog and replay-verified. See docs/concepts/determinism.md.

7.5 Capability outcomes are inputs

A capability call’s result is an input to the deterministic computation, not a deterministic output of it. Replay re-uses the recorded outcome rather than re-invoking the capability. This is what allows recording a run against a live network or LLM and replaying it offline with byte-identical results.

8. Replay-verified vs. operational state

Every byte of state the VM tracks falls into one of two classes. Replay verification covers only the first.

8.1 Replay-verified state

These are deterministic functions of the module + input + recorded event log. Replay MUST produce the same bytes:

  • The module on disk (constants, code, types, functions, globals, entry).
  • The operand stack at every instruction boundary, for the entire run.
  • The local frames of every active call.
  • The values pushed to / popped from the operand stack by each opcode.
  • Every Value materialized during execution.
  • The EventLog: SchedulerTick, ActorSpawn, MessageSend, MessageReceive, CapCall (request + recorded result).
  • The serialized output (Halt’s top-of-stack).
  • The evidence-bundle JSON serialization of the above.
  • Map iteration order (always key-sorted; §7.3).
  • Float arithmetic results (IEEE 754; §7.2).

8.2 Operational state

These exist in a real run but are NOT replay-verified. They MAY differ between record and replay:

  • Wall-clock time of execution (the time.now capability result is replay-verified; the timestamp on the bundle envelope is not).
  • VM step-count budgets used by execute_bounded (the schedule’s logical order is preserved; the per-step CPU budget is not part of the trace).
  • Host-process resource usage (RAM, file descriptors, network sockets).
  • The ui.render sink’s actual rendered pixels (EmitUi emits the tree; the host’s interpretation is operational).
  • Logger output, metrics, and Prometheus counters.

Cross-reference: docs/concepts/determinism.md for the rationale and worked examples; docs/concepts/capabilities.md for the capability-gating model.

9. Diagnostics (informative)

The reference implementation surfaces module-level errors via BytecodeError:

  • InvalidMagic — the first four bytes did not match MAGIC.
  • UnsupportedVersion(v) — the header version field was not 1.
  • InvalidBytecode(msg) — payload truncation, malformed code, or other structural defects.
  • Serialization(msg) — the JSON payload failed to encode/decode.

VM-level errors (stack underflow, type mismatch, capability denied, division by zero, list out-of-bounds, mod by zero, etc.) are surfaced through the VM’s VmError type; conformant implementations are not required to use the same names but SHOULD distinguish stack/type/capability/arithmetic categories.

10. Reader contract for 1.x VMs

A VM advertising 1.x conformance MUST:

  • Accept any 1.0 module. Forward-compat for additive changes within 1.x is the contract; if a 1.0 module presents itself, the VM MUST execute it without modification.
  • Reject bytecode_version >= 2. A typed UnsupportedVersion (or equivalent) MUST be returned. Silent acceptance of a future major is forbidden — this is the §1 application of “reject at parse, don’t silently override”.
  • Reject unknown opcodes. Any byte tag not in §4.2 (or §4.4’s reserved space, when the VM is 1.0-only) MUST trigger a typed unknown-opcode error rather than be silently treated as Nop.
  • Reject unknown capability IDs. A CapCall with cap_id outside the frozen 1.0 set (and not registered for the VM’s specific 1.y minor version) MUST fail with a typed capability-unknown error.
  • Preserve the determinism contract (§7) for all conformant inputs.

A 1.y VM (y > 0) accepting an additively-extended module MUST still reject any 2.x module and any opcode/capability not in its known set.

11. Cross-references

12. Change log for this specification

  • 1.0 (2026-04-28) — Initial freeze. Sprint W9-A. Captures the bytecode format as shipped in Boruna v1.0.0-rc2: magic LLMB, internal version 1, JSON-payload module wire format, 48 frozen opcodes, 15 Value variants, 11 capabilities at contract version "1", key-sorted Map iteration, deterministic actor scheduling.
  • 1.1 (2026-05-20) — Additive opcode minor bump per §1.2(6). Adds Op::Debug (0xA7) and Op::DebugMsg (0xA8) for __builtin_debug(v) / __builtin_debug_msg(msg, v) — stack-identity print-and-passthrough helpers writing to host stderr. Operational-only (§8.2); replay-verified state and capability gating unchanged. A 1.0 reader presented with either MUST reject with an unknown-opcode typed error (§10). Rationale in claudedocs/research_quint_borrowable_ideas_2026-05-20.md and docs/retro-quint-borrow-2026-05-20.md.

Workflow DAG


schema_version: 1 status: stable last_revised: 2026-04-28 sprint: W4

Workflow DAG Schema — version 1.0

This document is the formal contract for workflow.json files consumed by boruna workflow validate, boruna workflow run, the coordinator/worker HTTP protocol, and any downstream reader. It freezes the on-disk shape at sprint W4 of the 0.5.0 spec-freeze track.

Versioning

The schema follows a simple-semver rule:

  • The schema_version field is a single integer representing the current major version. Today: 1.
  • Minor revisions (“1.y”) add optional, additive fields. Readers built against any 1.x MUST accept any 1.y document (y >= x) and silently ignore fields they do not know.
  • Major revisions (“2.x”) signal a breaking change. Readers built for 1.x MUST refuse a schema_version: 2 document — they cannot guarantee correctness against a format they have not been taught.

This implies a forward-compatibility promise within a major and explicit operator action across a major.

Reader gate

The reference Rust reader is boruna_orchestrator::workflow::definition::WorkflowDef. It enforces the gate at parse time, per project conventions §1 (reject at parse, not later):

ConditionResult
schema_version field absentWorkflowParseError::MissingSchemaVersion
schema_version > 1WorkflowParseError::UnsupportedSchemaVersion { found, supported_max }
schema_version not an integerWorkflowParseError::InvalidJson
schema_version: 1 + valid bodyOk(WorkflowDef)

Stable surface strings (project conventions §2):

  • workflow.missing_schema_version
  • workflow.unsupported_schema_version
  • workflow.invalid_json

The constant boruna_orchestrator::WORKFLOW_DAG_SCHEMA_VERSION = 1 is the authoritative max-supported version for this build.

JSON Schema

Authoritative JSON Schema (Draft 2020-12) for a workflow.json document conforming to schema_version 1:

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://boruna.dev/spec/workflow-dag/1.0.json",
  "title": "Boruna Workflow DAG (schema_version 1)",
  "type": "object",
  "required": ["schema_version", "name", "version", "steps", "edges"],
  "properties": {
    "schema_version": {
      "description": "Major schema version. MUST be 1 for this spec.",
      "type": "integer",
      "minimum": 1,
      "maximum": 1
    },
    "name": {
      "description": "Human-readable workflow identifier.",
      "type": "string",
      "minLength": 1
    },
    "version": {
      "description": "Operator-controlled workflow version (semver string).",
      "type": "string",
      "minLength": 1
    },
    "description": {
      "description": "Optional free-form description.",
      "type": "string"
    },
    "steps": {
      "description": "Map of step id -> step definition. Order is canonicalized via BTreeMap on read.",
      "type": "object",
      "minProperties": 1,
      "additionalProperties": { "$ref": "#/$defs/Step" }
    },
    "edges": {
      "description": "Explicit DAG edges as [from_step_id, to_step_id] pairs. Edges may also be implied via a step's `depends_on`. Cycles are rejected by the validator.",
      "type": "array",
      "items": {
        "type": "array",
        "minItems": 2,
        "maxItems": 2,
        "items": { "type": "string" }
      }
    }
  },
  "$defs": {
    "Step": {
      "type": "object",
      "required": ["kind"],
      "properties": {
        "kind": {
          "description": "One of: \"source\", \"approval_gate\", \"external_trigger\".",
          "type": "string",
          "enum": ["source", "approval_gate", "external_trigger"]
        },
        "source": {
          "description": "Path (relative to workflow dir) to the .ax file. Required for kind=\"source\".",
          "type": "string"
        },
        "required_role": {
          "description": "Required reviewer role. Required for kind=\"approval_gate\".",
          "type": "string"
        },
        "condition": {
          "description": "Optional gate condition expression (kind=\"approval_gate\"). Informational; not currently enforced by the runner.",
          "type": ["string", "null"]
        },
        "confidence_gate": {
          "description": "Optional calibrated auto-approval (kind=\"approval_gate\"). Additive in 1.x. When the score clears the threshold the gate completes without a human; otherwise it pauses as usual. See docs/design-conformal-gating.md.",
          "type": "object",
          "required": ["source_step", "calibration", "alpha_permille"],
          "properties": {
            "source_step": {
              "description": "A source step in the gate's depends_on (or with an edge into the gate). Its `result` must be an Int in 0..=1000 (permille); anything else escalates to a human.",
              "type": "string"
            },
            "calibration": {
              "description": "Calibration file, relative to the workflow directory, no `..`. JSON: {\"version\":1,\"examples\":[{\"score\":0..1000,\"correct\":bool}]}.",
              "type": "string"
            },
            "alpha_permille": {
              "description": "Largest allowed chance that a wrong answer skips a human, in permille.",
              "type": "integer",
              "minimum": 1,
              "maximum": 999
            }
          }
        },
        "description": {
          "description": "Optional human-readable description (kind=\"external_trigger\").",
          "type": ["string", "null"]
        },
        "capabilities": {
          "description": "Capabilities this step is allowed to invoke. Subset of the workflow policy's allow-list.",
          "type": "array",
          "items": { "type": "string" },
          "default": []
        },
        "inputs": {
          "description": "Map of local input name -> upstream reference of the form \"<step_id>.<output_name>\". Cross-step data flow.",
          "type": "object",
          "additionalProperties": { "type": "string" },
          "default": {}
        },
        "outputs": {
          "description": "Map of output name -> declared type label (informational). The runner stores the actual step value under the key \"result\".",
          "type": "object",
          "additionalProperties": { "type": "string" },
          "default": {}
        },
        "depends_on": {
          "description": "Implicit edges: this step runs only after every listed step has completed. Combined with `edges` to form the full DAG.",
          "type": "array",
          "items": { "type": "string" },
          "default": []
        },
        "timeout_ms": {
          "description": "Wall-clock budget per attempt. Exceeding fails the attempt with `wall_time_exceeded`.",
          "type": ["integer", "null"],
          "minimum": 0
        },
        "retry": { "$ref": "#/$defs/RetryPolicy" },
        "budget": { "$ref": "#/$defs/StepBudget" }
      }
    },
    "RetryPolicy": {
      "type": ["object", "null"],
      "required": ["max_attempts"],
      "properties": {
        "max_attempts": {
          "description": "Upper bound on attempts. <= 1 means single-attempt always.",
          "type": "integer",
          "minimum": 1
        },
        "on_transient": {
          "description": "Legacy gate (used when `retry_on` is empty): true retries on any failure, false is single-attempt.",
          "type": "boolean",
          "default": false
        },
        "retry_on": {
          "description": "Allowlist of error classes that should trigger a retry. Unknown class strings are silently ignored (typo-safe; conservative-by-default).",
          "type": "array",
          "items": { "type": "string" },
          "default": []
        }
      }
    },
    "StepBudget": {
      "type": ["object", "null"],
      "properties": {
        "max_tokens": { "type": ["integer", "null"], "minimum": 0 },
        "max_calls":  { "type": ["integer", "null"], "minimum": 0 }
      }
    }
  }
}

Notes:

  • The schema does NOT use "additionalProperties": false. Forward- compatibility (§Versioning above) requires that 1.x readers accept additive fields introduced in 1.y minors.
  • The schema does NOT require description, capabilities, inputs, outputs, depends_on, timeout_ms, retry, budget. The reference reader applies serde defaults.
  • The schema does NOT enforce DAG acyclicity, edge endpoint existence, or input reference well-formedness — these are enforced by the validator (WorkflowValidator::validate), one layer above structural parsing.

Field-level documentation

Top-level

  • schema_version (required, integer = 1) — major schema version. See §Versioning and §Reader gate above.
  • name (required, non-empty string) — workflow identifier surfaced in CLI output, logs, and evidence bundles.
  • version (required, non-empty string) — operator-controlled workflow version (typically semver). Independent of schema_version; this versions the workflow, not the schema.
  • description (optional string) — free-form prose.
  • steps (required, non-empty map) — step_id -> step body. Step ids are unique within a workflow.
  • edges (required array) — [from_step_id, to_step_id] pairs. Combined with each step’s depends_on to form the DAG.

Step body

A step’s kind discriminates which fields are required:

  • kind: "source" — compile and run an .ax file. Requires source (path relative to the workflow directory). The step’s return value is stored as the canonical result output and is available to downstream steps via <step_id>.result.
  • kind: "approval_gate" — pause the run until an operator records an approval/rejection via boruna workflow approve. Requires required_role. Optional condition is informational. Optional confidence_gate lets the gate complete without a human when a calibrated confidence score is high enough (see docs/design-conformal-gating.md). It is not supported with --submit-only or the coordinator, which reject it at submit time.
  • kind: "external_trigger" — pause the run until an external event arrives via boruna workflow trigger <run-id> <step-id>. Optional description is operator-facing only.

Common optional fields (all kinds):

Backwards-compatibility commitment

For any reader built against schema_version 1.x:

  1. Within a major (1.x → 1.y, y > x): the reader MUST accept the document. Future minors only ADD fields; they MUST NOT change the meaning of existing fields, MUST NOT introduce new required fields, and MUST NOT tighten existing field types.
  2. Across a major (1.x → 2.0): the reader MUST refuse with a typed UnsupportedSchemaVersion error. The operator must upgrade Boruna or downgrade the workflow.

This commitment is exercised in unit tests at orchestrator/tests/workflow_integration.rs (sprint W4):

  • workflow_def_loads_with_schema_version_1
  • workflow_def_rejects_missing_schema_version
  • workflow_def_rejects_future_major_version
  • workflow_def_accepts_unknown_optional_fields (forward-compat)
  • example_workflows_all_validate

Replay invariant (§15)

schema_version is part of the canonical-JSON serialization that feeds WorkflowRunner::workflow_hash_from_def. Two workflows that differ only in schema_version produce different workflow_hash values, so evidence bundles correctly bind the schema generation under which a run was executed. Replay verification fails closed if the on-disk schema_version no longer matches the recorded hash.

Post-1.0 additive notes (no version bump)

These behaviors are additive extensions of the 1.0 spec — they are visible only to operators who opt in, do not change the on-disk shape of any 1.0 workflow, and do not change the schema_version: 1 constant.

  • Versioned worker capability advertisements (post1-T-1.3). RegisterRequest.advertised_capabilities (the wire shape used at worker registration, not part of the workflow def) now accepts {name, version} objects in addition to bare strings. The coord normalizes legacy strings to the coord’s current Capability::version() for that name. Routing and the new coord.capability_version_mismatch claim error are documented in docs/reference/error-kinds.md. No workflow JSON change.

Cross-references

Evidence Bundle Format Specification — version 1.0

Status: stable Format version: 1.0 Reader contract: semver-like — 1.x is forward-compatible for any 1.0 reader; 2.x is breaking and MUST be rejected.

This document is the source of truth for the on-disk layout, integrity contract, and version semantics of evidence bundles emitted by boruna workflow run --record and validated by boruna evidence verify.


1. Format-version semantics

Every bundle declares its format in a top-level bundle.json file. Readers MUST gate on this field before reading any other content.

Bundle format_version1.0 reader1.x reader (x ≥ 1)2.x reader
missing / pre-1.0reject (legacy)reject (legacy)reject
"1.0"acceptacceptreject
"1.5"accept (forward-compat)acceptreject
"2.0"reject (incompatible major)rejectaccept

Compatibility rules:

  • Same major: unknown fields are ignored. A 1.0 reader presented with a 1.5 bundle MUST treat it as valid and silently drop fields it doesn’t recognize.
  • Different major: the wire format is allowed to change in incompatible ways. Readers MUST refuse to interpret content from a major they don’t know.
  • Missing bundle.json: the bundle is pre-1.0 (legacy). Readers MUST reject and surface a hint pointing the user at boruna migrate evidence-bundle (planned, sprint W5-C).

This applies §1 of the project conventions: “reject at parse, don’t silently override”. A reader that silently accepts an unknown major would let a future bundle’s content be misinterpreted as the format the reader expects.

2. bundle.json schema

{
  "format_version": "1.0",
  "boruna_version": "0.6.0",
  "created_at": "2026-04-28T14:32:11.420Z",
  "run_id": "run-2026-04-28-abc123",
  "workflow_hash": "9c8a1f...",
  "components": [
    "audit_log.json",
    "env_fingerprint.json",
    "manifest.json",
    "outputs/",
    "policy.json",
    "workflow.json"
  ]
}

Field semantics:

FieldTypeRequiredDescription
format_versionstringyesSemver-like format version. The major component is the compat gate.
boruna_versionstringyesVersion of the Boruna binary that emitted the bundle (CARGO_PKG_VERSION). Diagnostic, not a compat gate.
created_atstring (RFC3339, UTC)yesWall-clock time the bundle was finalized.
run_idstringyesWorkflow run identifier, matches manifest.run_id.
workflow_hashstring (hex)yesSHA-256 of the workflow definition JSON, matches manifest.workflow_hash.
componentsstring[]yesSorted list of component file/directory names actually present in this bundle. Trailing / denotes a directory. Diagnostic; readers MUST NOT rely on this list to drive parsing.
encryptionobjectnoAdditive in 1.x (sprint W6-B). When present in manifest.json, the bundle’s file contents are AES-256-GCM envelope-encrypted. See §9. Absent → plaintext bundle (original 1.0 behavior).

Future minor versions (1.x) may add optional fields. Readers MUST tolerate them.

3. Bundle directory layout (1.0)

<bundle-dir>/
├── bundle.json             # version gate (this spec)
├── manifest.json           # cryptographic manifest with file checksums + bundle_hash
├── workflow.json           # snapshot of the workflow definition
├── policy.json             # snapshot of the active policy
├── audit_log.json          # hash-chained event log
├── env_fingerprint.json    # OS / arch / boruna_version captured at run time
└── outputs/
    └── <step_id>/
        └── <output_name>.json   # per-step JSON outputs (compact form)

bundle.json is the LAST file written during finalize. Every other component is written and (where applicable) parent-dir fsynced before bundle.json is committed via the same atomic-rename + parent-dir-fsync pattern used by workflow::data_flow::DataStore::store_output. Consequence: a reader observing bundle.json is guaranteed to observe a complete bundle.

3.1 Component contracts

ComponentContract
manifest.jsonBundleManifest (see orchestrator/src/audit/evidence.rs). Carries file_checksums: BTreeMap<filename, sha256> for every other file (excluding bundle.json and manifest.json itself).
workflow.jsonThe workflow definition as submitted, byte-for-byte. workflow_hash = sha256(workflow.json).
policy.jsonThe policy snapshot. policy_hash = sha256(policy.json).
audit_log.jsonAuditLog JSON; chain integrity is independently verifiable via AuditLog::verify.
env_fingerprint.jsonOS / arch / CARGO_PKG_VERSION of the recording binary.
outputs/<step>/<name>.jsonCompact JSON; same bytes that DataStore::hash_value hashed and that the orchestrator’s SQLite checkpoint persisted. sha256sum MUST match the output_hash recorded in the audit log.

4. Hash-chain integrity contract

Independent of the format gate, verify_bundle enforces:

  1. Every entry in manifest.file_checksums matches the SHA-256 of the on-disk file.
  2. audit_log.json parses as a valid AuditLog, and every entry’s entry_hash is sha256(prev_hash || event_json). The chain is broken iff any entry fails this check.
  3. audit_log.hash() (last entry’s entry_hash) equals manifest.audit_log_hash.
  4. All required components from §3 are present.

A bundle that fails any of (1)–(4) is INVALID. The verify_bundle reader emits the failing checks but does not attempt to “repair” — that is reserved for boruna migrate.

5. Reader compat matrix

Feature1.0 reader1.x reader (x ≥ 1)
Read 1.0 bundleyesyes
Read 1.x bundle (x > 0)yes (drops unknown fields)yes
Read pre-1.0 / legacy bundleno (reject + migration hint)no
Read 2.x bundlenono

Producers MUST emit the lowest format version that contains every field they need; this maximizes the population of readers that can consume the bundle.

6. Future evolution (non-normative)

The 0.5-S7 retro flagged the sidecar layout for output blob references as the leading driver of a future 1.1 minor bump. Rationale: large LLM step outputs are stored content-addressed (/api/runs/{id}/blobs/{hash}) since 0.5-S7. Bundles currently inline-resolve the bytes into outputs/<step>/result.json. A 1.1 bundle would optionally carry a blobs/<sha256> sidecar directory and rewrite outputs/<step>/result.json to a { "$blob_ref": "<hash>" } reference. A 1.0 reader would (correctly) reject the rewritten outputs JSON unless it understands $blob_ref; therefore the sidecar layout is a 1.x field-level extension rather than a parse-time break — provided producers continue to emit the inline form when no blobs are referenced. The compat matrix is preserved.

A 2.0 break — for example, switching from JSON to a binary-framed format — is not currently planned and would require an ADR plus a migration path through boruna migrate evidence-bundle.

7. Migration

Bundles produced by Boruna v0.5.0 and earlier do NOT carry bundle.json. The reader rejects them with:

unsupported evidence bundle format_version: found `missing bundle.json (legacy bundle from pre-1.0 release; use `boruna migrate evidence-bundle` to upgrade)`, expected major `1`

The boruna migrate evidence-bundle tool is planned for sprint W5-C. Until it ships, legacy bundles must be re-recorded against a current binary.

8. Encryption (additive 1.x)

The optional encryption field in manifest.json carries the metadata for AES-256-GCM envelope encryption (sprint W6-B). When absent, the bundle is plaintext (the original 1.0 behavior). When present, file contents in the bundle directory are AES-256-GCM ciphertext keyed off a fresh per-bundle data-encryption key (DEK) which is itself wrapped with an operator-supplied key-encryption-key (KEK).

Field shape:

FieldTypeRequiredDescription
algorithmstringyesLocked to "aes-256-gcm" for 1.x
kek_idstringyesOperator-supplied identifier; bound as AAD into the wrapped DEK
wrapped_dekbase64yesDEK encrypted with the KEK
wrapped_dek_noncebase64yesNonce for the DEK wrap
filesstring[]noOperational hint listing encrypted files; readers MUST NOT use this to drive parsing

algorithm, kek_id, wrapped_dek, and wrapped_dek_nonce are replay-verified — they live inside manifest.json and therefore participate in bundle_hash. files is OPERATIONAL metadata (per project convention §15): it is informational; the canonical set of encrypted files is derivable from manifest.file_checksums and tampering with files cannot bypass decryption (the verify loop iterates file_checksums, not encryption.files).

Per-file nonces are deterministic SHA-256(filename)[..12]. This is safe because the DEK is freshly generated per bundle and filenames are unique within a bundle, so a (DEK, nonce) pair never repeats. Each per-file ciphertext carries an AES-GCM authentication tag; tag failure on read surfaces as evidence.cipher_tag_invalid.

8.1 Reader contract for encryption

A 1.x reader:

  • MUST accept bundles WITHOUT encryption (plaintext path; original 1.0 behavior).
  • MUST accept bundles WITH encryption if a KEK is available.
  • MUST reject encryption.algorithm values other than "aes-256-gcm" with evidence.unsupported_algorithm.
  • MUST surface DEK-unwrap failure as evidence.encryption_key_mismatch and a missing KEK as evidence.encryption_key_required.
  • MUST surface per-file tag mismatch as evidence.cipher_tag_invalid and refuse to return decrypted bytes from the failing file.
  • MUST NOT log the unwrapped DEK at any severity.

8.2 Reject-at-parse contract (§1)

Per project convention §1 (“reject at parse, don’t silently override”), a reader MUST refuse to interpret a bundle when:

  • encryption is present but algorithm is anything other than "aes-256-gcm" → evidence.unsupported_algorithm.
  • encryption is present but a required field (kek_id, wrapped_dek, wrapped_dek_nonce) is missing or malformed (non-base64, wrong length) → reader-defined parse error.
  • encryption is present but no KEK has been supplied to the reader → evidence.encryption_key_required.

A 1.0 reader (pre-W6-B) that does not know the encryption field will ignore it as an unknown field per §1 of this spec, then fail integrity verification because the on-disk bytes hash to ciphertext rather than plaintext referenced by manifest.file_checksums. Operators using a 1.0 reader against an encrypted bundle MUST upgrade to a 1.x reader that understands encryption. This is the documented compat story per §B.3 of the LTS contract (docs/lts.md).

9. References

  • Implementation: orchestrator/src/audit/evidence.rs (BundleJson, EvidenceBundleBuilder::finalize)
  • Reader gate: orchestrator/src/audit/verify.rs (check_bundle_format, verify_bundle)
  • Constant: orchestrator/src/audit/mod.rs (BUNDLE_FORMAT_VERSION)
  • Concept doc: docs/concepts/evidence-bundles.md
  • CLI surface: boruna evidence verify, boruna evidence inspect [--json]

Boruna Runtime-Provenance Predicate 1.0

predicateType: https://boruna.dev/runtime-provenance/v1

This spec defines an interop view of a Boruna evidence bundle: the same provenance the native bundle already records (manifest.json), re-emitted as a standard in-toto Statement wrapped in a DSSE envelope so that off-the-shelf supply-chain tooling (cosign verify-blob, in-toto-verify) can consume Boruna evidence.

It is additive and non-breaking. Producing an attestation does not modify the bundle, its manifest.json, its bundle_hash, or the existing evidence verify path. See docs/spec/evidence-bundle-1.0.md for the native format that remains the source of truth.

The producer/verifier lives in orchestrator/src/audit/attestation.rs; the CLI surface is boruna evidence attest.


1. Artifacts

Two nested artifacts are produced from a finalized BundleManifest:

  1. an in-toto Statement v1, and
  2. a DSSE envelope wrapping that Statement, signed with the SAME ed25519 key used for manifest signing (EvidenceBundleBuilder::with_signing_key). No new keypair is introduced; the DSSE keyid is the lowercase-hex ed25519 public key, identical to manifest.signature.public_key.

The envelope is written to <bundle-dir>/attestation.intoto.dsse.json.


2. Statement shape

{
  "_type": "https://in-toto.io/Statement/v1",
  "subject": [
    { "name": "<component-file-path>", "digest": { "sha256": "<hex>" } },
    // ... one per manifest.file_checksums entry ...
    { "name": "boruna-bundle:<run_id>", "digest": { "sha256": "<bundle_hash>" } }
  ],
  "predicateType": "https://boruna.dev/runtime-provenance/v1",
  "predicate": { /* see §3 */ }
}

Subjects

subject[] is the set of artifacts this attestation makes claims about:

  • Every component file the manifest checksums — reusing manifest.file_checksums verbatim (name → sha256). These include workflow.json, policy.json, audit_log.json, env_fingerprint.json, per-step outputs/<step>/<name>.json, and any optional components (intents.json, model_invoking_steps.json, event_log.json).
  • The bundle itself, as a synthetic subject boruna-bundle:<run_id> whose sha256 is the manifest’s bundle_hash. (bundle_hash is the SHA-256 of the canonicalized manifest, so it is a digest, not a file — but it lets a verifier bind the whole bundle by a single value.)

file_checksums is a BTreeMap, so subjects are emitted in sorted-name order — the Statement bytes are deterministic for a given manifest.


3. Predicate schema

The predicate borrows the shape of SLSA Provenance v1’s buildDefinition / runDetails, mapping the existing manifest fields:

{
  "buildDefinition": {
    "buildType": "https://boruna.dev/workflow-run/v1",
    "externalParameters": {
      "workflowName": "<manifest.workflow_name>",
      "workflowHash": "<manifest.workflow_hash>",
      "policyHash":   "<manifest.policy_hash>"
    },
    "internalParameters": {
      "borunaVersion": "<producing boruna version>",
      "envFingerprint": { /* manifest.env_fingerprint, verbatim */ }
    }
  },
  "runDetails": {
    "builder": { "id": "https://boruna.dev/boruna@<version>" },
    "metadata": {
      "invocationId": "<manifest.run_id>",
      "startedOn":    "<manifest.started_at>",
      "finishedOn":   "<manifest.completed_at>"
    },
    "byproducts": {
      "auditLogHash": "<manifest.audit_log_hash>",
      "bundleHash":   "<manifest.bundle_hash>"
    }
  }
}

Field mapping (manifest → predicate)

Manifest fieldPredicate location
workflow_namebuildDefinition.externalParameters.workflowName
workflow_hashbuildDefinition.externalParameters.workflowHash
policy_hashbuildDefinition.externalParameters.policyHash
env_fingerprintbuildDefinition.internalParameters.envFingerprint
(build version)buildDefinition.internalParameters.borunaVersion
run_idrunDetails.metadata.invocationId
started_atrunDetails.metadata.startedOn
completed_atrunDetails.metadata.finishedOn
audit_log_hashrunDetails.byproducts.auditLogHash
bundle_hashrunDetails.byproducts.bundleHash

Capability set and contract-check results are not manifest struct fields — they live in bundle components (event_log.json carries ContractCheck events; model_invoking_steps.json lists steps that reached an llm.* capability). Those components appear as subjects (by SHA-256) rather than being inlined into the predicate, so the attestation still binds them without duplicating or re-parsing their contents. A consumer that wants contract-check detail dereferences the event_log.json subject and reads it from the bundle.

All maps are BTreeMap and struct fields serialize in declaration order, so the Statement is byte-stable: same run → same manifest → same Statement bytes → same signature.


4. DSSE envelope

{
  "payload": "<base64(canonical Statement JSON bytes)>",
  "payloadType": "application/vnd.in-toto+json",
  "signatures": [
    { "sig": "<base64(ed25519 sig over PAE)>", "keyid": "<hex ed25519 pubkey>" }
  ]
}

The signature is an ed25519 signature over the DSSE Pre-Authentication Encoding (PAE) of (payloadType, payload):

PAE(type, body) = "DSSEv1" SP LEN(type) SP type SP LEN(body) SP body

where SP is a single ASCII space (0x20) and LEN is the ASCII-decimal byte length. type is application/vnd.in-toto+json and body is the raw Statement JSON bytes (the pre-base64 bytes), not the base64 text.

Known vector (from the DSSE spec’s worked example, asserted in attestation.rs tests):

PAE("http://example.com/HelloWorld", "hello world")
  = "DSSEv1 29 http://example.com/HelloWorld 11 hello world"

The payloadType is bound into the signed bytes, so a signature over one payload type cannot be replayed under another.


5. CLI

# Produce: sign the manifest's provenance into a DSSE envelope.
#   Reuses the SAME ed25519 seed you signed the bundle with.
boruna evidence attest <bundle-dir> --signing-key <64-hex-seed>
#   (or set BORUNA_BUNDLE_SIGNING_KEY instead of --signing-key)
# → writes <bundle-dir>/attestation.intoto.dsse.json

# Verify: check the DSSE signature over the PAE.
boruna evidence attest <bundle-dir> --verify
# Optionally pin the trusted signer key:
boruna evidence attest <bundle-dir> --verify --verify-key <64-hex-pubkey>

--verify exits non-zero on any failure (bad signature, mutated payload, wrong payloadType, or a pinned key that made no valid signature).


6. Ecosystem compatibility

The envelope is a standard DSSE envelope with payloadType application/vnd.in-toto+json and a standard in-toto Statement payload, so it is structurally consumable by the wider ecosystem:

  • cosign verify-blob-attestation consumes a DSSE envelope + a public key and verifies the ed25519 signature over the PAE. Export the signer’s public key in PEM form (the DSSE keyid here is the raw 32-byte ed25519 public key as hex; cosign expects a PEM PUBLIC KEY, so wrap the key in SubjectPublicKeyInfo DER → PEM before handing it to cosign). The signature algorithm (ed25519 over PAE) and envelope layout match what cosign verifies.
  • in-toto-verify / the in-toto attestation validators parse the Statement (_type, subject, predicateType, predicate) directly; the predicate is a custom type, so policy is expressed against predicateType == https://boruna.dev/runtime-provenance/v1 and the fields in §3.

Honest caveat. The bytes and algorithms follow the DSSE and in-toto specs, and Boruna verifies its own envelopes end-to-end (round-trip, PAE known-vector, and tamper tests). Full black-box interop with a specific cosign / in-toto release — including the exact public-key PEM/DER encoding each tool wants and any tool-specific envelope expectations — has not been exercised against those binaries in this change; treat cross-tool verification as “spec-conformant, pending a live cosign/in-toto-verify integration check.” The key-encoding bridge (raw hex ed25519 → SPKI PEM) is the most likely point of friction.

Boruna Application Framework Specification

Overview

The framework defines a mandatory application protocol for all Boruna programs. Every application follows a strict structure: init → update → view → effects cycle.

The VM is the kernel. The framework is userland. The runtime does not depend on the framework.

1. Application Protocol

Every app must implement four functions:

fn init() -> State
fn update(state: State, msg: Message) -> UpdateResult
fn view(state: State) -> UITree
fn policies() -> PolicySet

Rules

  • update() must be pure — no capability annotations allowed.
  • update() returns UpdateResult { state: State, effects: List<Effect> }.
  • view() must be pure — returns a declarative UITree.
  • init() may use capabilities for initial setup.
  • policies() declares required capabilities and constraints.

Compile-Time Validation

The framework compiler validates:

  • All four functions exist with correct signatures.
  • update() has no capability annotations.
  • view() has no capability annotations.
  • State type is defined and serializable.
  • Message type is an enum.

2. Effect System

Effects are declarative descriptions of side effects:

type Effect {
    kind: String,       // one of the built-in effect kinds below
    payload: Value,     // structured payload (type depends on effect kind)
    callback_tag: String,  // message tag for delivering the result
}

Built-in effect kinds:

  • http_request — maps to net.fetch capability
  • db_query — maps to db.query capability
  • fs_read — maps to fs.read capability
  • fs_write — maps to fs.write capability
  • timer — maps to time.now capability
  • random — maps to random capability
  • spawn_actor — creates child actor
  • emit_ui — emits UI tree to host

The framework runtime executes effects between update cycles. Effect results are delivered as messages to the next update() call.

3. State Management

  • State must be a record type.
  • State is serialized to JSON between cycles for snapshots.
  • Framework provides:
    • snapshot(state) — serialize state to JSON string
    • restore(json) — deserialize state from JSON string
    • diff(old, new) — produce list of changed fields

4. UI Model

type UINode {
    tag: String,
    props: String,
    children_json: String,
}

UITree is a UINode at the root, with children encoded as JSON. The view function returns a UINode.

Constraints:

  • Pure function of State.
  • No side effects.
  • Host renders the tree.
  • User events become Messages fed to update().

5. Actor Integration

  • Child actors use the same App protocol.
  • Parent spawns child via spawn_actor effect.
  • Messages between actors are routed by the framework runtime.
  • Supervision: if a child crashes, parent receives an error message.

6. Policy Layer

type PolicySet {
    capabilities: List<String>,
    max_effects_per_cycle: Int,
    max_steps: Int,
}

Policy violations:

  • Abort safely with structured error.
  • Error is replay-compatible.

7. Testing Harness

Built-in testing functions:

  • simulate(init, messages) — run message sequence, return final state
  • assert_state(state, field, expected) — check state field
  • assert_effects(effects, expected_kinds) — check effect kinds
  • replay_verify(log1, log2) — compare execution logs

Testing does not require a host UI.

8. Implementation

The framework is a Rust crate boruna-framework that provides:

  • AppValidator — compile-time validation of App protocol
  • AppRuntime — execution loop for the App protocol
  • EffectExecutor — maps effects to capability calls
  • StateMachine — state transition engine with snapshot/diff
  • TestHarness — testing utilities

The framework compiles .ax sources through the normal compiler, then wraps execution in the App protocol runtime.

Multi-Agent Orchestration Layer — Specification

1. Overview

The orchestrator enables parallel, safe, deterministic development on the Boruna codebase. It is a separate tool — it does not modify the compiler, VM, or framework runtime. It coordinates work by scheduling tasks as a DAG, enforcing a two-person rule (Implementer + Reviewer), managing file-level locks to prevent conflicts, and gating all changes through deterministic CI checks.

2. Work Graph Model

2.1 Nodes

Each node represents a unit of work:

WorkNode {
    id: String,              // unique identifier (e.g., "WN-001")
    description: String,     // human-readable summary
    inputs: Vec<String>,     // file paths or artifact IDs consumed
    outputs: Vec<String>,    // file paths or artifact IDs produced
    dependencies: Vec<String>, // IDs of nodes that must complete first
    owner_role: Role,        // Planner | Implementer | Reviewer
    tags: Vec<String>,       // e.g., ["compiler", "vm", "framework"]
    status: NodeStatus,
    assigned_to: Option<String>,
    patch_bundle: Option<String>, // path to .patchbundle.json
    review_result: Option<ReviewResult>,
}

2.2 Node States

pending  → ready → running → passed
                  ↘ blocked
                  ↘ failed
StateMeaning
pendingDependencies not yet met
readyAll dependencies passed; eligible for assignment
runningAssigned to a role; work in progress
blockedLock conflict or external dependency
failedGate check failed; may retry
passedAll gates passed; work accepted

2.3 Edges

Directed edges encode dependency: A → B means B cannot start until A is passed. The graph must be a DAG (no cycles). The engine validates acyclicity on plan creation.

2.4 Scheduling

The engine uses topological sort to determine execution order:

  1. Compute in-degree for each node.
  2. Enqueue nodes with in-degree 0.
  3. When a node passes, decrement successors’ in-degree.
  4. Nodes reaching in-degree 0 become ready.

Concurrency is bounded by max_parallel (default: 4). Retry policy: transient failures (exit code > 128) retry up to 2 times with 1s delay. Permanent failures (exit code 1) do not retry.

3. Roles

RoleResponsibility
PlannerCreates the work graph (DAG), assigns tags, defines dependencies
ImplementerProduces patch bundles for assigned nodes
ReviewerReviews patch bundles, runs gate checks, approves or rejects
Red-team (optional)Adversarial review — tries to break the change

3.1 Two-Person Rule

Every node that modifies code requires two steps:

  1. Implement: An Implementer produces a patch bundle and marks the node running.
  2. Review: A Reviewer runs boruna-orch review <bundle>, which:
    • Validates bundle format
    • Runs deterministic gates (compile, test, replay)
    • Checks reviewer checklist items
    • Outputs approve or reject

A node can only reach passed if both steps succeed. The Implementer and Reviewer must be different agents (enforced by assigned_to field).

4. Artifact Types

4.1 Patch Bundle (.patchbundle.json)

{
  "version": 1,
  "metadata": {
    "id": "PB-20260220-001",
    "intent": "Add list_set opcode for indexed mutation",
    "author": "agent-1",
    "timestamp": "2026-02-20T10:30:00Z",
    "touched_modules": ["boruna-bytecode", "boruna"],
    "risk_level": "low"
  },
  "patches": [
    {
      "file": "crates/boruna-bytecode/src/opcode.rs",
      "hunks": [
        {
          "start_line": 45,
          "old_text": "    ListPush,          // 0x83",
          "new_text": "    ListPush,          // 0x83\n    ListSet,           // 0x87"
        }
      ]
    }
  ],
  "expected_checks": {
    "compile": true,
    "test": true,
    "replay": true,
    "diagnostics_count": null
  },
  "reviewer_checklist": [
    "No new language features introduced",
    "Backward compatible with existing bytecode",
    "Tests cover happy path and error cases"
  ]
}

4.2 Diagnostics Report

Structured JSON output from boruna-orch report --json:

{
  "graph_id": "G-001",
  "total_nodes": 5,
  "passed": 3,
  "failed": 1,
  "pending": 1,
  "nodes": [...],
  "locks": [...],
  "last_gate_results": {...}
}

4.3 Trace Hash

A stable hash produced by boruna framework trace-hash used to verify determinism. Stored per-node as part of gate results.

5. Conflict Rules

5.1 Module-Level Locking

For MVP, locking operates at the crate/module level:

Lock TargetGranularityExample
CrateEntire crate directorycrates/boruna-bytecode
ExampleExample directoryexamples/admin_crud
DocSingle filedocs/language-guide.md

5.2 Lock Lifecycle

  1. When a node transitions to running, locks are acquired for all outputs.
  2. If any lock is held by another node, the requesting node becomes blocked.
  3. Locks are released when the node reaches passed or failed.
  4. Stale locks (node stuck in running > timeout) can be force-released via boruna-orch status --force-unlock <node-id>.

5.3 Conflict Detection

Before applying a patch bundle:

  1. Check that no locked module overlaps with touched_modules.
  2. Verify file checksums match expected state (patches apply cleanly).
  3. If conflict detected, the apply operation fails and the node becomes blocked.

6. Deterministic Gating Process

Every patch bundle must pass these gates in order:

GateCommandPass Criteria
1. Compilecargo build --workspaceExit code 0
2. Testcargo test --workspaceAll tests pass
3. Replaycargo run -- framework trace-hash <file>Hash matches expected
4. Lint (optional)cargo clippy --workspaceNo errors

Gate results are recorded per-node:

{
  "node_id": "WN-001",
  "gates": {
    "compile": { "status": "pass", "duration_ms": 3200 },
    "test": { "status": "pass", "duration_ms": 8100, "total": 179, "passed": 179 },
    "replay": { "status": "pass", "hash": "a1b2c3d4e5f6g7h8" }
  }
}

If any gate fails, the node becomes failed and the patch bundle is rolled back (best-effort).

7. Storage

MVP uses local JSON files under orchestrator/storage/:

orchestrator/storage/
  graphs/
    G-001.json        # work graph
  bundles/
    PB-*.patchbundle.json
  locks/
    locks.json        # active lock table
  gates/
    WN-001.gate.json  # per-node gate results

8. CLI Commands

CommandDescription
boruna-orch plan <spec.json>Create DAG from a plan specification
boruna-orch next --role <role>Assign next ready node for the given role
boruna-orch apply <bundle.patchbundle.json>Apply patch bundle, run gates
boruna-orch review <bundle.patchbundle.json>Review bundle: validate + gates + checklist
boruna-orch statusShow current graph state
boruna-orch report --jsonMachine-readable summary of graph + gates

9. Adapter Interface

Adapters wrap existing tooling as “judges”:

#![allow(unused)]
fn main() {
trait GateAdapter {
    fn name(&self) -> &str;
    fn run(&self, context: &GateContext) -> GateResult;
}
}

Built-in adapters:

  • CompileAdapter — runs cargo build --workspace
  • TestAdapter — runs cargo test --workspace, parses test counts
  • ReplayAdapter — runs cargo run -- framework trace-hash, compares hashes
  • DiagAdapter — runs cargo run -- framework diag, captures JSON output

10. Non-Goals (MVP)

  • Network-distributed agents (local-only for MVP)
  • Git integration (manual commits; orchestrator doesn’t touch git)
  • AST-level patch granularity (file-level hunks for MVP)
  • Red-team role automation (manual for now)
  • Real-time collaboration (sequential role handoff)

Package Ecosystem Specification

Overview

Deterministic, content-addressed package system for the Boruna platform. No remote registries, no version ranges, no dynamic loading. All resolution is exact and reproducible.

Package Manifest (package.ax.json)

Every package has a manifest at its root:

{
  "name": "example.package",
  "version": "0.1.0",
  "description": "Short description",
  "dependencies": {
    "other.package": "0.2.1"
  },
  "required_capabilities": ["net.fetch", "db.query"],
  "exposed_modules": ["core", "utils"],
  "integrity": "sha256:<hex>"
}

Fields

FieldTypeRequiredDescription
namestringyesDotted package name (e.g. std.collections)
versionstringyesSemver (MAJOR.MINOR.PATCH)
descriptionstringyesHuman-readable description
dependenciesmapnoPackage name → exact version
required_capabilitiesarraynoCapability strings from boruna-bytecode::Capability
exposed_modulesarrayyesModule names this package exposes
integritystringcomputedContent hash, set by boruna-pkg publish

Validation Rules

  • name must match ^[a-z][a-z0-9]*(\.[a-z][a-z0-9]*)*$
  • version must be valid semver: MAJOR.MINOR.PATCH
  • dependencies must specify exact versions (no ranges, no wildcards)
  • required_capabilities must be valid capability names from boruna-bytecode::Capability
  • exposed_modules must contain at least one entry
  • integrity is computed at publish time; absent in development

Lockfile (ax.lock.json)

Generated only by the resolver. Never hand-edited.

{
  "lockfile_version": 1,
  "resolved": {
    "example.package@0.1.0": {
      "integrity": "sha256:abc123...",
      "dependencies": {
        "other.package": "0.2.1"
      }
    },
    "other.package@0.2.1": {
      "integrity": "sha256:def456...",
      "dependencies": {}
    }
  }
}

Rules

  • Generated by boruna-pkg resolve
  • Execution must use lockfile exclusively
  • Missing or invalid lockfile is a hard error at build/run time
  • Lockfile must be committed to version control
  • Changing lockfile requires reviewer approval (orchestrator gate)

Package Storage (Local Registry)

Content-addressed layout:

packages/
  registry/
    <package-name>/
      <version>/
        package.ax.json
        src/
          <module>.ax
        bytecode/
          <module>.axbc
        HASH

Content Hash

The HASH file contains sha256:<hex> computed over:

  1. All source files (src/**/*.ax) sorted by path
  2. The manifest (package.ax.json) with integrity field removed
  3. Dependency hashes (sorted by package name)

This produces a Merkle-like hash: changing any transitive dependency changes the root hash.

Publish Flow

  1. Validate manifest
  2. Compile all exposed modules to bytecode
  3. Compute content hash
  4. Set integrity field in manifest
  5. Copy to registry under <name>/<version>/

Resolver

Algorithm

  1. Parse root manifest dependencies
  2. For each dependency, load its manifest from registry
  3. Recursively resolve transitive dependencies
  4. Topological sort (Kahn’s algorithm)
  5. Detect circular dependencies → hard error
  6. Detect conflicts (two versions of same package) → hard error
  7. Generate lockfile with all resolved packages and their hashes

Invariants

  • Same manifest + same registry → identical lockfile (deterministic)
  • No version range resolution
  • No SAT solver
  • Fail fast on any ambiguity

Capability Enforcement

When building an application:

  1. Parse root manifest’s required_capabilities
  2. For each dependency (transitively), collect required_capabilities
  3. Union all capabilities
  4. Check against application policy
  5. If any dependency requires a forbidden capability → compile error

Policy Format

Application policy is specified in the root manifest or a separate policy.ax.json:

{
  "allowed_capabilities": ["net.fetch", "time.now"],
  "denied_capabilities": ["fs.write", "random"]
}

If allowed_capabilities is present, only those are permitted. If denied_capabilities is present, those are blocked. Cannot specify both. If neither, all capabilities are allowed.

CLI Commands

CommandDescription
boruna-pkg initCreate package.ax.json in current directory
boruna-pkg add <pkg> <version>Add dependency to manifest
boruna-pkg remove <pkg>Remove dependency from manifest
boruna-pkg resolveGenerate ax.lock.json from manifest + registry
boruna-pkg installResolve + verify all packages exist in registry
boruna-pkg publishCompile, hash, copy to local registry
boruna-pkg verifyVerify all installed packages match their hashes
boruna-pkg treePrint dependency tree

Orchestrator Integration

When a patch bundle touches package.ax.json or ax.lock.json:

  • Gate adapter re-runs boruna-pkg resolve and verifies lockfile matches
  • Lockfile changes require reviewer approval (two-person rule)
  • Capability changes in dependencies trigger policy re-evaluation

Security Model (MVP)

  • Local registry only (no remote fetch)
  • No auto-download
  • Manual publish/install
  • Hash verification mandatory on every install/build
  • No code execution during install

Non-Goals (MVP)

  • Remote package registry
  • Version range resolution
  • Optional dependencies
  • Platform-specific packages
  • Pre/post install scripts
  • Binary distribution

Determinism Contract

What Is Deterministic

Given the same bytecode module and the same message sequence:

  1. State transitions: Every update(state, msg) produces the identical Value output.
  2. Effect lists: The effects returned by update() are identical in kind, payload, and order.
  3. UI trees: The view(state) output is identical for identical state.
  4. Scheduling order: In single-actor mode, messages are processed FIFO.
  5. Cycle count: The number of cycles matches exactly.

What Must Be Virtualized

These sources of non-determinism are kept outside the pure core:

SourceHow Virtualized
Timetimer effect → capability gateway → logged in events
Randomnessrandom effect → capability gateway → logged in events
Network I/Ohttp_request effect → capability gateway → logged
File I/Ofs_read/fs_write effects → capability gateway
Databasedb_query effect → capability gateway

All external interactions go through the capability gateway, which logs every call and result in the EventLog. Replay substitutes recorded results, guaranteeing identical execution.

What Must Be Logged

The framework CycleRecord logs per cycle:

  • cycle — cycle number
  • message — the input message (tag + payload)
  • state_before — state value before update
  • state_after — state value after update
  • effects — effect list returned by update
  • ui_tree — view output

The VM EventLog logs:

  • CapCall — capability name + arguments
  • CapResult — capability name + return value
  • UiEmit — emitted UI tree
  • ActorSpawn, MessageSend, MessageReceive, SchedulerTick

Replay Contract

  1. Record: Run the app with a real capability handler. Save the EventLog.
  2. Replay: Run the same bytecode with ReplayHandler seeded from recorded CapResult values.
  3. Verify: The replay EventLog must produce identical CapCall sequences (same capability, same args, same order).

If verification fails, either:

  • The bytecode is non-deterministic (bug in compiler/VM).
  • An external value leaked into the pure core (bug in framework).

Multi-Actor Determinism

In single-actor mode, messages are processed FIFO. In multi-actor mode, the VM uses round-robin scheduling across actors. The exact scheduling order is captured in the EventLog via SchedulerTick, ActorSpawn, MessageSend, and MessageReceive events.

During replay, the EventLog enforces the identical scheduling sequence. This means multi-actor execution is deterministic as long as it is replayed from the same EventLog. Two independent runs with the same bytecode and messages are NOT guaranteed to produce the same scheduling order — only record-then-replay guarantees identical execution.

If your application requires fully deterministic multi-actor ordering without replay, restrict to single-actor mode or use explicit message sequencing in your update logic.

Enforcement

  • update() and view() run with Policy::deny_all() — no capability calls allowed.
  • Effects are the only way to request external data.
  • The capability gateway is not accessible during pure function execution.
  • Violation = FrameworkError::PurityViolation.

Golden Test Protocol

Golden tests hash the following after running a fixed message sequence:

  1. Final state JSON snapshot.
  2. Full cycle log (state_before, state_after, effects per cycle).
  3. Concatenated effect kind strings.

If the hash changes, the test fails with a diff showing exactly what diverged.

Boruna Enterprise Platform Overview

Boruna is a deterministic execution platform for enterprise AI workflows. It provides policy-gated, auditable workflow execution with built-in governance, replay, and compliance evidence generation.

Core Capabilities

Workflow Execution

  • DAG-based workflow definitions with typed data flow between steps
  • Topological execution ordering with dependency resolution
  • Approval gates for human-in-the-loop review
  • Retry policies for transient failures
  • Budget enforcement per step and per workflow

Determinism Guarantees

  • Same inputs + same workflow + same policy = identical outputs
  • BTreeMap-based ordering throughout (no HashMap non-determinism)
  • Capability-gated side effects — all IO is declared and controlled
  • Record/replay support via EventLog

Policy Enforcement

  • Capability allowlists: declare which side effects each step may use
  • Budget limits: token and call count budgets per step
  • Model allowlists: restrict which LLM models may be invoked
  • Network allowlists: restrict outbound HTTP destinations

Audit Trail

  • Hash-chained audit log: tamper-evident, append-only record of all decisions
  • Evidence bundles: self-contained compliance artifacts per workflow run
  • Bundle verification: cryptographic integrity checking of all artifacts

Architecture

workflow.json → Validator → Runner → Evidence Bundle
                  ↓           ↓
              Topological   Per-step:
              Sort          Compile .ax → VM → Output
                            Policy check
                            Audit log entry

Crate Map

  • boruna-orchestrator — Workflow engine, audit system, evidence bundles
  • boruna-compiler — Compiles .ax source to bytecode
  • boruna-vm — Executes bytecode with capability gating
  • boruna-bytecode — Bytecode format and Value types
  • boruna-effect — LLM integration with budget tracking
  • boruna-framework — App protocol (Elm architecture)
  • boruna-tooling — Diagnostics, repair, trace-to-tests, templates
  • boruna-pkg — Package system with integrity verification

Workflow Lifecycle

  1. Define — Write workflow.json with steps, edges, policies
  2. Validate — boruna workflow validate <dir> checks DAG structure
  3. Run — boruna workflow run <dir> --policy <policy> executes steps
  4. Record — --record flag generates evidence bundle
  5. Verify — boruna evidence verify <dir> checks bundle integrity
  6. Replay — Re-execute from recorded event log for determinism verification

Schema Versioning

All serializable formats include schema_version for forward compatibility:

  • Workflow definition: v1
  • Policy: v1
  • Audit log: v1
  • Evidence bundle manifest: v1

Security Model

Capability System

All side effects in Boruna are declared and enforced. Functions annotate their capabilities:

fn fetch(url: String) -> String !{net.fetch} { ... }

The VM’s CapabilityGateway checks every capability call against the active Policy. Undeclared capabilities are blocked at compile time; unauthorized capabilities are blocked at runtime.

Built-in Capabilities

  • net.fetch — HTTP requests
  • db.query — Database queries
  • fs.read, fs.write — File system access
  • llm.prompt — LLM model invocation
  • actor.spawn, actor.send — Actor system operations

Isolation

Process Isolation

Each workflow step compiles to bytecode and runs in a fresh VM instance. Steps cannot share memory or state except through the explicit data flow system.

Filesystem Isolation

  • PatchBundle validates against .. and absolute paths
  • canonicalize() defense-in-depth prevents path traversal
  • Evidence bundle outputs are written to controlled directories

Data Validation

  • LLM cache keys are hex-only validated
  • Context store hashes are hex-only validated
  • Package content uses SHA-256 integrity verification

Policy Enforcement

Policies can restrict:

  • Which capabilities a step may use
  • Which LLM models may be invoked
  • Which network endpoints are reachable
  • Token and call budgets per step

Policy violations are recorded in the audit log with deny decisions.

Secrets Management (Gap)

Currently there is no dedicated secrets management system. Secrets should NOT be passed as environment variables in production. This is documented as a P1 gap. The recommended interim approach:

  • Use external secret managers (Vault, AWS Secrets Manager)
  • Pass secrets via capability handlers at runtime
  • Never embed secrets in workflow definitions or source files

Threat Model

In Scope

  • Malicious workflow steps: Capability gating prevents unauthorized IO
  • Tampered evidence: Hash-chained audit log and bundle checksums detect modification
  • Path traversal: Validated at multiple layers
  • Non-determinism: BTreeMap ordering, controlled randomness, replay verification

Out of Scope (current)

  • Network-level attacks (TLS termination is external)
  • Host OS compromise
  • Supply chain attacks on the Boruna binary itself
  • Multi-tenant isolation (single-tenant model currently)

Audit Trail

Every policy decision, capability invocation, and approval action is recorded in the hash-chained audit log. The chain is cryptographically verifiable:

  • Each entry hashes: sequence number + previous hash + event data
  • Tampering with any entry breaks the chain from that point forward
  • boruna evidence verify checks chain integrity

Digital Signatures (Gap)

Evidence bundles use SHA-256 checksums but do not yet support Ed25519 digital signatures. This is documented as a P2 gap.

Platform Governance

Policy System

Boruna enforces policies at multiple levels:

VM-Level Policy

The Policy struct controls capability access at the VM level:

  • rules: BTreeMap<String, PolicyRule> — per-capability allow/deny with budget
  • default_allow: bool — default behavior for undeclared capabilities
  • schema_version: u32 — for forward compatibility

Built-in policies: Policy::allow_all(), Policy::deny_all().

Framework-Level PolicySet

For framework apps, PolicySet adds application-specific constraints:

  • max_cycles — maximum update cycles
  • allowed_effects — effect type allowlist
  • max_effects_per_cycle — rate limit on effects

LLM Policy

LlmPolicy controls LLM-specific behavior:

  • total_token_budget — total tokens allowed
  • max_context_tokens — per-request context limit
  • model_allowlist — approved model identifiers
  • cache_policy — caching behavior (Always, Never, Conditional)

Budget Enforcement

Budgets are enforced at the step level:

{
  "budget": {
    "max_tokens": 10000,
    "max_calls": 5
  }
}

When a budget is exceeded, the step fails with an auditable error.

Approval Gates

Workflow steps can be approval gates that pause execution:

{
  "kind": "approval_gate",
  "required_role": "reviewer",
  "condition": "severity >= 3"
}

When reached, the workflow pauses with status Paused and records an ApprovalRequested audit event.

Audit Log

Every workflow run produces a hash-chained audit log. Events include:

  • WorkflowStarted — workflow and policy hashes
  • StepStarted / StepCompleted / StepFailed
  • CapabilityInvoked — with allow/deny decision
  • PolicyEvaluated — rule evaluation details
  • BudgetConsumed — token/call consumption
  • ApprovalRequested / ApprovalGranted / ApprovalDenied
  • WorkflowCompleted — result hash and duration

Each entry’s hash includes the previous entry’s hash, forming a tamper-evident chain.

RBAC (Gap)

Currently, policies are per-run rather than per-user. A full RBAC system is documented as a P1 gap in ENTERPRISE_GAPS.md. The current model:

  • Workflow author defines the policy
  • CLI operator selects which policy to apply
  • Approval gates specify a required role (string-based, not yet identity-verified)

Performance Baseline

This document is the published baseline for Boruna’s performance budget (roadmap milestone for 1.0.0). It captures reproducible benchmarks, not portable absolute numbers — your machine will measure differently, and that’s fine. The goal is a stable harness so we can detect regressions over time.

Benchmark suite

The suite lives in benches/ at the workspace root and uses criterion 0.5. It covers three areas:

  • compile.rs — boruna_compiler::compile() end-to-end (lex → parse → typeck → codegen). Three sizes: ~20-line program, ~200-line program with records / pattern matching / loops, and the post-substitution crud-admin template.
  • vm_throughput.rs — Vm::run() on tight loops at 1k / 10k / 100k iterations. Variants: pure arithmetic, per-iteration record allocation, and a 4-deep call-chain stand-in for capability dispatch (the surface language only generates CapCall from step_input, which needs a workflow context — pure call dispatch shares the same hot opcode loop).
  • evidence.rs — EvidenceBundleBuilder::finalize() and verify_bundle() round-trips at 0 / 5 / 10 steps.

How to run

# Build the harness without running (fast, used in CI)
cargo bench -p boruna-benches --no-run

# Run all benches with criterion's defaults (~3 minutes)
cargo bench -p boruna-benches

# Run a single bench file with a quick configuration
cargo bench -p boruna-benches --bench compile -- \
    --sample-size 10 --warm-up-time 1 --measurement-time 3

Criterion writes detailed JSON + (optionally) HTML reports under target/criterion/. Compare two runs with criterion --baseline <name> after saving a baseline (--save-baseline <name>).

Baseline (recorded 2026-04-28)

Captured with --sample-size 10 --warm-up-time 1 --measurement-time 3 on a developer laptop (macOS, x86_64, dev profile of dependencies but release profile for the bench binary — criterion’s default). The median column is what we compare against; the bracket shows criterion’s lo/hi bound for the central tendency.

BenchmarkMedianRange
compile_small_program32 µs26 – 44 µs
compile_medium_program346 µs273 – 511 µs
compile_crud_admin_template397 µs318 – 452 µs
vm_pure_loop/iters=1000361 µs332 – 397 µs
vm_pure_loop/iters=100003.97 ms3.58 – 4.66 ms
vm_pure_loop/iters=10000036.0 ms32.2 – 41.6 ms
vm_record_loop/iters=10001.47 ms1.29 – 1.84 ms
vm_record_loop/iters=1000032.5 ms27.5 – 36.2 ms
vm_call_dispatch_loop/iters=10005.42 ms3.80 – 6.62 ms
vm_call_dispatch_loop/iters=1000037.7 ms32.2 – 46.1 ms
evidence_build_empty2.16 ms1.93 – 2.67 ms
evidence_build_5_steps7.22 ms6.37 – 8.44 ms
evidence_verify_5_steps3.95 ms3.23 – 4.65 ms
evidence_verify_10_steps5.71 ms3.76 – 7.48 ms

Roughly: a small program compiles in tens of microseconds, a medium program in hundreds of microseconds; the VM sustains ~2.7 M arithmetic-loop iterations per second; an evidence bundle round-trip (build + verify, 5 steps) takes ~11 ms total.

1.x performance budget commitments

These are conservative ceilings — roughly 2x the baseline median plus headroom — that we commit to NOT regressing past on hardware in the same class as the baseline machine. CI does not enforce them today (see “CI” below); they’re a contract for human review of perf- sensitive PRs.

BudgetCeilingSource bench
compile_small_program< 5 mscompile_small_program (median 32 µs)
compile_medium_program< 5 mscompile_medium_program (median 346 µs)
compile_crud_admin_template< 5 mscompile_crud_admin_template (median 397 µs)
vm_pure_loop per 100k iters< 100 msvm_pure_loop/iters=100000 (median 36 ms)
vm_record_loop per 10k iters< 80 msvm_record_loop/iters=10000 (median 32 ms)
evidence_build_5_steps< 25 msevidence_build_5_steps (median 7.2 ms)
evidence_verify_10_steps< 50 msevidence_verify_10_steps (median 5.7 ms)

If a PR drops a number more than 2x past the ceiling, treat it as a regression — bisect, profile, and either fix or document the cause before merging.

Interpreting regressions

  1. Save a baseline before changing anything:
    cargo bench -p boruna-benches -- --save-baseline before
    
  2. Apply your change.
  3. Compare:
    cargo bench -p boruna-benches -- --baseline before
    
    Criterion prints a per-benchmark verdict (Improved, Regressed, No change).
  4. If Regressed shows up:
    • Re-run on a quiet machine (close browsers, stop background processes — criterion is sensitive to scheduling jitter).
    • If still regressed, look at flamegraphs (cargo flamegraph --bench <name>) before guessing.
    • A 5–10 % wobble between runs is normal; flag only consistent regressions of 25 %+ relative to the budget headroom.

CI

Benches are not gated in CI today — cargo bench is slow (each benchmark needs ≥3 s of measurement time × ≥10 samples for stable numbers) and criterion’s stdout is noisy. The smoke test in benches/tests/smoke.rs runs under cargo test --workspace and guarantees the bench fixtures still compile and execute, which catches the most common breakage (a lib refactor stranding a bench call).

A future sprint may add a perf-comparison gate via criterion-compare-action or cargo-codspeed, both of which run on hosted runners with bare-metal stability and post a PR comment with the diff. The decision is gated on actual CI flake rate — regression detection is worse than no detection if random green builds get reported as 30 % regressions.

For now, the workflow is: run locally before merging perf-sensitive PRs.

Source layout

benches/
  Cargo.toml          # workspace member, depends on criterion 0.5
  src/lib.rs          # shared fixtures (.ax sources, bundle builder)
  benches/
    compile.rs        # compile time benchmarks
    vm_throughput.rs  # vm step throughput benchmarks
    evidence.rs       # evidence bundle write/verify benchmarks
  tests/
    smoke.rs          # CI-gated smoke test exercising every fixture

Sprint reference: W5-A (1.0.0 roadmap entry “Performance benchmarks”).

Roadmap

This roadmap describes what Boruna is working toward. It is realistic, not aspirational marketing. Items without a milestone are under consideration but not scheduled.

Last refreshed: 2026-05-17 (after the v1.4.0 release).

v3.0.0 — HTTP / distributed layer removed

As of v3.0.0, Boruna’s entire HTTP / serving / distributed-execution layer has been removed: the distributed coordinator, distributed workers, active-active HA, and coordinator mTLS; the three web UIs (workflow dashboard, evidence web viewer, approval console); the serve cargo feature and its server dependencies; and the coordinator, dashboard, worker, and evidence serve CLI commands plus the --coordinator / --coord-token flags.

Boruna is now a local deterministic engine + CLI. The historical 0.4.0 / 0.5.0 milestones below record distributed-execution work that shipped at the time and has since been removed — they are retained as history, not as descriptions of current capabilities. Approval and external-trigger gates remain, handled locally via boruna workflow approve/reject/trigger plus resume.

Current: 1.4.0 — SHIPPED (2026-05-17)

Workspace version is 1.4.0. Fourth feature minor on the 1.x LTS line. Agent-native CLI inspection surfaces (boruna doctor, boruna size, boruna workflow graph, boruna lang codes, boruna skills — all --json-capable, motivated by a competitive review of vercel-labs/zero); the boruna-lsp language server for .ax files (diagnostics, completion, formatting); three compliance example workflows (SOC 2 audit, HIPAA data pipeline, financial review). See the CHANGELOG for the full list.

Previous: 1.3.0 — SHIPPED (2026-04-30)

Workspace version was 1.3.0. Third feature minor on the 1.x LTS line. 27 new __builtin_* functions (string, list, and map operations); import resolution wired end-to-end (import "std-name" inlines libs/<name>/src/core.ax at compile time); evidence inspect now shows step output content for plaintext bundles (500-char preview in text mode, full "step_outputs" in --json); std-llm and std-json graduated to 1.0-stable (all 13 stdlib packages are now stable); std-json gains json_array, fixed int_to_string and json_escape; std-validation gains validate_contains, validate_starts_with, validate_ends_with. See the CHANGELOG for the full list.

Previous: 1.2.0 — SHIPPED (2026-04-29)

Workspace version was 1.2.0. Second feature minor on the 1.x LTS line. All 11 original std-* packages graduated to 1.0-stable; improved error messages and lang repair; compliance templates, model eval, and LSP MVP landed as future-track work. See the CHANGELOG for the full list.

Previous: 0.3.0

Released 2026-04-26 — closes every big-rock theme on the original 0.3.0 plan: persistent workflow state (crash-resumable), concurrent step execution within waves, step retry policies, idempotent invocation, workflow versioning for CI/CD safety, the LLM-handler decision (BYOH), per-step attempt tracking with the project’s first schema migration, workflow step output piping via the step_input builtin, and async step execution via the external-trigger CLI for webhook-driven workflows. Plus review-driven safety work (atomic trigger-commit closing a TOCTOU race; SSRF-hardened real HTTP handler).

See CHANGELOG.md for the full 0.3-S2a → 0.3-S16 sprint stack.

Previous: 0.2.0

Released 2026-04-25 — driven by FleetQ implementer feedback. Closes the two P0 adoption blockers; other P1/P2 asks tracked as issues #3–#9.

What shipped:

  • Fine-grained capability policy in MCP boruna_run — accepts a structured Policy object (per-capability allow/budget rules, allowlist vs. denylist mode, NetPolicy with allowed_domains / methods / byte limits / timeout), in addition to the legacy "allow-all" / "deny-all" strings. Documented JSON Schema 2020-12 at docs/reference/policy.schema.json. Breaking (MCP only): unknown policy values now return error_kind: "invalid_policy" instead of silently treating them as "allow-all".
  • Multi-target static binary releases on every v* tag: x86_64-unknown-linux-musl, aarch64-unknown-linux-musl, x86_64-apple-darwin, aarch64-apple-darwin, plus combined SHA256SUMS. Linux builds are musl so they run on Alpine and other libc-minimal distros.
  • docs/releasing.md — release process and verification.

What did NOT ship from the original 0.2.0 plan (deferred to 0.2.x or 0.3.0):

  • boruna new interactive scaffold
  • boruna fmt auto-formatter
  • Watch mode (boruna run --watch)
  • Improved error messages with suggested fixes for all common mistakes
  • Better lang repair coverage
  • Evidence bundle diff
  • Workflow step output piping
  • std-llm, std-json library expansion

These were displaced by the FleetQ adoption work. They are still on the path to v1 — see 0.2.x and 0.3.0 below.

0.2.x — Developer experience patch lane

Target: rolling, May–July 2026

The DX work originally scoped for 0.2.0 ships incrementally as point releases. Each is small, additive, and non-breaking.

  • boruna new — scaffold a new workflow from a template interactively (sprint W3-C)
  • boruna fmt — auto-formatter for .ax files (v1: post1-T-1.3; v2 comment-preserving: post1/fmt-v2 PR #45)
  • boruna run --watch — re-run on file change (post1-T-1.4)
  • Improved error messages — boruna lang check now suggests nearest variable/function name for E003/E004 errors (edit-distance-1), type-conversion hints for E009, improved message text for E001/E002/E007 (post1/improved-diagnostics-repair)
  • Better lang repair — handles E003/E004 near-miss rename patches; bottom-up patch ordering prevents offset corruption; new Conservative strategy applies only high-confidence (≥80%) patches (post1/improved-diagnostics-repair)
  • Evidence bundle diff — boruna evidence diff <bundle-a> <bundle-b> (post1/evidence-diff, PR #44)
  • Expanded stdlib — std-llm, std-json libraries (post1/std-new-packages PR #46)

0.3.0 — Real-use durability — SHIPPED (2026-04-26)

Focus: workflows that survive process restarts, handle long-running steps, and unblock production use cases. Combined original 0.3.0 plan with two FleetQ P1 asks that fit thematically.

  • Persistent workflow state — checkpoint and resume across process restarts (0.3-S2a/S2b/S3/S6)
  • Async step execution — steps that wait for external events via webhook-driven CLI trigger (0.3-S15); approval gates (0.3-S2c)
  • Step retry policies — configurable retry with backoff on transient failures (0.3-S5)
  • Workflow versioning — --expect-workflow-hash for CI/CD safety (0.3-S9)
  • Workflow step output piping — step_input builtin (0.3-S14)
  • Structured resource limits with typed errors (#5) — max_memory_mb, max_wall_ms, max_output_bytes (0.3-S10)
  • Versioned capability identity (#3) — boruna_capability_list returns capability_set_hash for safe caching
  • LLM live handler decision — DECIDED (sprint 0.3-S8): Bring Your Own Handler (BYOH). No default LLM handler ships in core; integrators wire their provider via the CapabilityHandler trait. Rationale + integration contract + reference OpenAI handler in docs/guides/llm-integration.md.
  • Concurrent step execution within waves — --concurrency N (0.3-S4)
  • Idempotent invocation — --skip-if-running for cron-driven scheduling (0.3-S7/S10)
  • Per-step attempt tracking with first schema migration v1→v2 (0.3-S11/S12/S13)
  • Atomic trigger commit closing TOCTOU race (0.3-S16, review-driven)
  • Scheduled workflows — trigger workflows on a cron schedule (deferred to 0.3.x; partially addressed by --skip-if-running for safe cron invocation) — boruna workflow schedule <dir> --cron "..." cron daemon (post1/scheduler-registry-rolling)

0.4.0 — Operations (mostly shipped on master, tag pending)

Originally targeted Q4 2026; landed early as 0.4-S1 → 0.4-S16 + 0.5-S1 → 0.5-S2f. Tag will be cut after auth (0.5-S3) lands.

  • Distributed step execution — coord+workers HTTP cluster + multi-wave advancement (0.5-S2a → 0.5-S2f)
  • Workflow dashboard — Axum + askama SSR (0.4-S16); merged into the coordinator listener (0.5-S2d)
  • Prometheus metrics endpoint — /metrics route + per-run-status counters (0.4-S?)
  • OpenTelemetry observability (#9) — per-capability OTLP spans (0.4-S5)
  • Policy management as code — Policy JSON files + boruna policy validate (0.4-S?)
  • Multi-environment support — --env flag + namespaced data-dir + Prometheus env= label (0.4-S14)
  • Streaming output from boruna_run (#4) — periodic progress events + capability call markers (post1-T-1.1, post1-T-2.2)
  • LLM provider registry — config-driven provider selection via --providers providers.json; ProviderRegistry validates config + logs intent (post1/scheduler-registry-rolling)
  • Scheduled workflows (carried over from 0.3.x) — full cron daemon via boruna workflow schedule (post1/scheduler-registry-rolling)

0.5.0 — Distributed execution + spec freeze

Target: ~Q3-Q4 2026 (accelerated from original Q1 2027 target).

Two sub-themes: (a) finish what 0.5-S2* started so distributed mode is production-grade, (b) lock the API surface for 1.0.

(a) Distributed-execution closure

  • workflow run --submit-only + coordinator wait — end-to-end multi-wave (0.5-S2e/f)
  • 0.5-S3 — Authentication — shared-secret bearer token. MUST land before any non-loopback bind is recommended. Gating for production deployments.
  • W6-A — mTLS + per-worker client certificates — additive opt-in mTLS surface on the coord HTTP routes. Cert subject CN drives worker identity; mismatch returns coord.identity_mismatch. Bearer auth path remains unchanged for LTS compatibility. See docs/design-coord-mtls.md.
  • 0.5-S4 — workflow run --coordinator <url> — combines submit + wait in one command for CI workflows
  • 0.5-S5 — Distributed retry policies — wires RetryPolicy through the wait driver so failed steps with retry budget transition Failed → Pending instead of permanent Failed
  • 0.5-S6 — Distributed approval-gate / external-trigger — generalizes the operator-bridge protocol from 0.3-S15 to work in distributed mode
  • 0.5-S7 — Output blob references — large step outputs (>64 KiB) stored in content-addressed blob store; inline/blob routing in runner; BlobStore read-side restore (post1/output-blob-refs)
  • Coordinator HA / failover (sprint W2) — multi-coord active-active against shared SQLite, worker URL failover at registration, /api/health for LB probes. The ADR 002 “coord restart = all leases void” assumption was audited and confirmed already-safe (threshold-based sweep preserves healthy leases under concurrent coords).
  • Worker capability tagging / placement (sprint W3-A) — workers advertise a SUBSET of the coord’s capability set via --advertise-caps; coord filters claims to caps the worker covers. Backwards-compatible (omitted flag = full fleet). New coord.unknown_capability error_kind.
  • Blob GC sweep (sprint W3-B) — boruna evidence gc-blobs reclaims orphan blobs in <data-dir>/blobs/. Closes the 0.5-S7 accepted limitation around manual cleanup.
  • Rolling upgrades — per-capability version negotiation via --advertise-cap-versions cap=ver; coordinator filters by version compatibility (post1/scheduler-registry-rolling)

(b) Spec freeze

  • Stable, documented MCP tool response schemas (#6) — protocol_version: 1 (0.5-S4 of FleetQ track)
  • Output JSON Schema validation as first-class gate (#8) (0.5-S6 of FleetQ track)
  • Record/replay for net.fetch (#7) (0.5-S7 of FleetQ track)
  • Versioned .ax language specification — formal grammar, type rules, capability semantics. Each future release publishes against a language_version. (Sprint W1-B, docs/spec/ax-language-1.0.md, boruna_compiler::LANGUAGE_VERSION = "1.0".)
  • Versioned workflow DAG schema — JSON Schema for workflow.json with schema_version field; backwards-compatible parser. (sprint W4; spec at docs/spec/workflow-dag-1.0.md, boruna_orchestrator::WORKFLOW_DAG_SCHEMA_VERSION = 1.)
  • Versioned evidence bundle format — schema for the bundle directory contents, format_version field, forward-compat reader. Shipped sprint W1-C; spec at docs/spec/evidence-bundle-1.0.md.
  • Versioned bytecode format — opcode discriminants, value model, capability ID table, module wire format, determinism contract. Shipped sprint W9-A; spec at docs/spec/bytecode-1.0.md, boruna_bytecode::BYTECODE_VERSION = "1.0".
  • Migration tooling beta — boruna migrate <from-version> upgrade path for any pre-1.0 breaking change. (sprint W5-C)

1.0.0 — Production readiness — SHIPPED (2026-04-28)

Milestone: the stable API surface is locked. 0.5+ programs compile and run unchanged. This is mostly a commitment release, not a feature release — the engineering between 0.5 and 1.0 is small but the durability promise is large.

  • Security audit of the VM and capability enforcement (external auditor; bookable months in advance — must commit Q4 2026 to land Q2 2027)
  • Performance benchmarks — published baseline for compile time, step throughput, evidence bundle write/verify time (sprint W5-A; see PERFORMANCE.md)
  • Long-term support commitment for 1.x — backports for security fixes, deprecation policy (sprint W5-B; see lts.md)
  • Migration tooling — boruna migrate covering pre-1.0 breaking changes (sprint W5-C)
  • All schemas (language, DAG, evidence, bytecode) finalized and documented
  • Evidence bundle encryption — at-rest encryption for bundles containing sensitive data (sprint W6-B, AES-256-GCM envelope encryption; see docs/design-bundle-encryption.md)

Previous: 1.1.0 — SHIPPED (2026-04-29)

First minor release on the 1.x LTS line. All changes are additive — no breaking changes.

  • MCP streaming capability call markers — boruna_run progress notifications carry "cap: llm.call" or "caps: llm.call, net.fetch" when capability calls fire during a VM slice. Gives MCP clients real-time visibility into what the VM is executing (post1-T-2.2).
  • Evidence bundle web inspector — boruna evidence serve <bundle-dir> [--port N] opens a local axum HTTP server with bundle overview, hash-chained audit log, and per-step output accordion. Verification runs inline. Feature-gated (boruna-cli/serve). Experimental tier (post1-T-4.4).
  • Trivia-in-AST foundation — new lex_full(source) API returns tokens with leading_trivia (attached // comments). Foundation for the comment-preserving boruna fmt v2 formatter. Existing lex() is unchanged. Experimental tier (post1-T-2.5).
  • BYOH reference handler library — four new CapabilityHandler reference implementations in examples/llm_handlers/: Anthropic Messages API, Ollama, vLLM/OpenAI-compatible, AWS Bedrock skeleton. Each is ~80–120 LOC, copy-and-tweak, no Cargo dep (post1-T-1.2).
  • BundleStorage adapters stable — S3, GCS, and Azure Blob adapters promoted from #[doc(hidden)] to stable public API. StorageError marked #[non_exhaustive]. New boruna evidence rotate-kek command re-encrypts DEKs under a new key-encryption key without touching ciphertext (post1-T-3.1–3.3, T-4.3).

What we need to decide now (before 0.3.0 starts)

These decisions block downstream planning. None of them are urgent today, but each one becomes urgent within 1–2 quarters.

  1. Security audit booking — pick auditor, scope, budget by Q4 2026. A real audit costs $30–100k and books months in advance. If this slips past Q4 2026, v1.0.0 slips with it.
  2. LLM live handler shipping plan — decided (0.3-S8): Bring Your Own Handler. See docs/guides/llm-integration.md.
  3. Persistence storage backend — decided (ADR 001): sqlite, no abstraction trait. Shipped via 0.3-S2a/S2b/S3/S6.
  4. Dashboard scope and tech — full SSR Rust stack (Axum + askama, fits the project) vs. SPA (more work, more polish). 0.4.0 dashboard depends on this answer.

Future / under consideration

These items are on the long-term radar but not scheduled:

  • Commercial platform: hosted workflow execution, managed evidence storage, SSO, RBAC, compliance reporting — built on the open source core.
  • IDE integration: language server (LSP) for .ax syntax, completion, and diagnostics in VS Code / Neovim (boruna-lsp MVP: diagnostics, completion, formatting — future/lsp).
  • Model evaluation framework: run the same workflow against multiple LLM providers and compare evidence bundles. (boruna workflow eval — future/model-eval)
  • Compliance templates: pre-built workflow patterns for common regulated use cases (SOC 2, HIPAA, financial audit — future/compliance-templates).
  • Cross-language FFI: call into Rust/Python libraries from .ax through a typed capability interface.

What is intentionally out of scope

Boruna will not become:

  • A general-purpose programming language (use Rust, Python, etc. for that)
  • An LLM framework (use LangChain, LCEL, etc. for that)
  • A cloud provider (Boruna runs where you deploy it)
  • A no-code tool (Boruna is for engineers)

Tracking

  • Filed issues: https://github.com/escapeboy/boruna/issues
  • FleetQ feedback issues: #3 through #9
  • Past sprint retros: retro/

See also: Stability, Limitations, Releasing

Security Policy

Supported Versions

Only the current release receives security patches.

VersionSupported
0.1.x✓

Once 1.0 ships, the long-term-support contract in docs/lts.md takes effect: 1.x is supported actively for 18 months from 1.0 GA and receives security fixes for 24 months. The 0.x line is EOL on 1.0 GA.

Reporting a Vulnerability

Do not open a public GitHub issue for security vulnerabilities.

Report using GitHub Security Advisories.

Include in your report:

  • Description of the vulnerability
  • Steps to reproduce
  • Potential impact and affected versions

Response Timeline

  • Acknowledgment: within 48 hours
  • Initial triage: within 5 business days
  • Status updates: every 7 days until resolved
  • Target resolution: within 90 days for critical issues

Scope

In scope: boruna-vm (capability gateway, replay engine), boruna-compiler, boruna-orchestrator (workflow runner, evidence bundle verification), boruna-mcp server.

Out of scope: example files, documentation, third-party dependencies.

Backport Policy

Security fixes are backported to every supported 1.y minor line for which the vulnerability applies. Fix versions are cut as patch releases (e.g. 1.3.4) on each affected line; patch releases contain the security fix and any trivially related test or doc changes only.

Severity follows CVSS v4:

  • CRITICAL or HIGH — fix released within 7 days of confirmed disclosure (or an interim advisory with mitigations if no fix is ready).
  • MEDIUM — fix released within 30 days of confirmed disclosure.
  • LOW — bundled with the next scheduled patch release on each supported line.

Pre-1.0, only the latest 0.x release receives fixes. Full backport contract and support-window definitions live in docs/lts.md.

Disclosure Policy

We follow coordinated disclosure:

  1. Reporter submits privately via GitHub Security Advisories
  2. We confirm the issue and assess severity
  3. We develop and test a fix
  4. We release the fix and publish a security advisory
  5. Reporter is credited (unless they prefer anonymity)

Safe Harbor

Good-faith security research conducted in accordance with this policy constitutes authorized testing. We will not pursue legal action for responsible vulnerability disclosure.

Changelog

All notable changes to Boruna are documented here. Format follows Keep a Changelog. Versioning follows Semantic Versioning.

Unreleased

[3.5.0] — 2026-10-03

Install in one command on every supported platform, three match bugs fixed, and a language specification that matches what the compiler accepts. Additive for programs that compiled before, but see Fixed: programs that use integer patterns or nested match now behave correctly, so their bytecode and module hashes change.

Added

  • install.sh (Linux, macOS) and install.ps1 (Windows): one-line install of the release binaries. They pick the build for the OS and CPU, verify it against SHA256SUMS and install nothing on a mismatch. Tested in CI on all five native platforms, including a tampered archive and Windows PowerShell 5.1.
  • Diagnostic E010 (warning): a binding declared without mut, a parameter or a for loop variable is reassigned. The compiler has accepted this since let mut was introduced, because the mut flag was parsed but never checked. boruna lang check and the MCP boruna_check tool now report it, and boruna lang repair adds the missing mut. It stays a warning in language 1.x and becomes an error in language version 2.0.

Changed

  • Language version is now 1.1. The specification (docs/spec/ax-language-1.0.md) now covers let mut, assignment, while and for (§4.5), which the compiler has accepted since v2.0 while the spec still listed them as reserved words. No program that compiled before stops compiling.
  • std-guard and std-json declare their loop counters with let mut (found by E010).

Fixed

  • match on a Some(x), Ok(x) or Err(x) value written in source never took the Some/Ok/Err arm (“no match found for value”); only values returned by builtins matched. The literals now build the same values builtins return, so they also compare equal to them.
  • Integer literal patterns (match n { 3 => ... }) never matched; only _ did. They now compile to equality checks, like string patterns.
  • A match inside an arm of another match could run the outer match against the inner arms and return a wrong result or fail.
  • A string or integer match with no matching arm and no _ arm returned () or failed later with “stack underflow”; it now fails with “no match found for value”, like other matches.
  • Because of these fixes, the bytecode of programs that use integer patterns or nested match changes, and so do their module hashes. A step that returned a Some/Ok/Err literal now records it as Some(..)/Ok(..)/Err(..) instead of an internal enum value.
  • docs/limitations.md said .ax has no mutable variables and no loops; it has both.

[3.4.0] — 2026-10-03

Platform release. Boruna now ships and is tested natively on macOS (Apple Silicon and Intel), Windows (x64 and Arm) and Linux (x86_64 and arm64). No change to workflows, workflow hashes or evidence bundles.

Added

  • --version / -V on all four binaries (boruna, boruna-mcp, boruna-pkg, boruna-orch).
  • Release binaries for macOS Intel (x86_64-apple-darwin) and Windows (x86_64-pc-windows-msvc, aarch64-pc-windows-msvc, packaged as .zip). Previously the README listed macOS Intel but no such binary was published, and there was no Windows build.
  • CI job that builds, tests and runs an example workflow with evidence verify natively on macOS arm64, macOS Intel, Windows x64, Windows on Arm and Linux arm64.
  • .gitattributes pins LF line endings so golden files and the changelog are identical on Windows.

Fixed

  • Windows: boruna overflowed the 1 MB main-thread stack on every command except --version. It now runs on a 64 MB stack thread on all platforms.
  • Windows: .ax sources with CRLF line endings failed to compile (unexpected character "\r").
  • Windows: patch bundles with a rooted path such as /etc/passwd passed the absolute-path check, because such a path has no drive letter. Rooted, drive-letter and backslash-rooted paths are now rejected on every platform.
  • Workflow trigger tokens read /dev/urandom and so could not be created on Windows; they now use the operating system’s random source and fail instead of falling back to weak randomness.

[3.3.0] — 2026-10-03

Additive feature release — no breaking changes. Two ideas from the agentlanguages.dev review: a calibrated way to let confident answers skip a human reviewer, and agent documentation that is generated from the binary so it cannot drift from the commands that exist. Existing workflows, workflow hashes and evidence bundles are unchanged.

Added

  • Calibrated confidence for approval gates — borrowed from conformal prediction (Quasar; the agentlanguages.dev review). An approval_gate can carry a confidence_gate (source_step, calibration, alpha_permille). When the upstream step’s Int score (permille) reaches the threshold computed from the calibration file, the gate completes without a human. Otherwise it pauses as before. The guarantee: a wrong answer is auto-approved with probability at most alpha. Too few wrong calibration examples means the threshold is never. The decision, score, threshold and the exact calibration file go into the evidence bundle (confidence_gates.json, confidence/<step>.calibration.json) and the hash-chained audit log; boruna evidence verify recomputes every decision, so a changed decision, a swapped calibration or a dropped file is rejected. Gate decisions cannot be redacted (evidence redact refuses them and verify rejects a redacted policy entry), because a redacted entry could hide one. The limits of verification are in docs/architecture-conformal-gating.md. evidence report --framework eu-ai-act lists auto-approved gates and marks Art. 14 as PARTIAL, since a gate that completed on a score is not human oversight. boruna confidence threshold shows the threshold before wiring it up. Works in-process (sequential, concurrent waves, resume). --submit-only and the coordinator reject workflows that use it. Existing workflow hashes are unchanged. See docs/design-conformal-gating.md and examples/workflows/confidence_gated_review.
  • Agent docs that cannot drift — borrowed from Vow and Lume (agentlanguages.dev review). boruna skills get cli now ends with a command reference generated from the installed binary. Half of the top-level commands (12 of 24) were absent from the hand-written text. boruna skills emit <dir> writes the skills as SKILL.md folders; boruna skills pack "<query>" --budget N returns only the relevant sections, deterministically.
  • A unit test fails when a hand-written skill names a boruna <command> that does not exist.

Fixed

  • ax-language skill pointed agents at boruna check, which does not exist. The command is boruna lang check.
  • crates/llmvm/src/actor.rs: drain(..).collect() replaced with std::mem::take (rust 1.98 clippy drain_collect, which -D warnings rejects). Behavior is unchanged.

[3.2.0] — 2026-07-18

Additive feature release — no breaking changes. Closes the two credibility gaps in Boruna’s tamper-evidence story identified during the adjacent-market research: external witnessing (so the record isn’t just “trust the recorder”) and privacy (so a sealed bundle can carry redactable data). Together they compose — a redaction preserves audit_log_hash, so an out-of-band anchor distinguishes an authorized redaction from a tamper.

Added

  • Transparency-log anchoring — boruna evidence anchor <dir> submits the bundle’s attestation to a Sigstore Rekor log and stores the inclusion proof + signed entry timestamp back into the bundle, giving an external witness and a trusted timestamp. --rekor-url points at a private Rekor for air-gapped use; --offline emits the entry payload for out-of-band submission; --verify checks a stored proof (RFC 6962 Merkle inclusion math) with no network. Network is opt-in, behind the rekor cargo feature (ureq); the default build stays network-free. Keyless (Fulcio) signing is design-noted (orchestrator/docs/keyless-signing.md).
  • Verifiable redaction — boruna evidence redact <dir> --event <i> [--field <f>] removes PII from a sealed bundle without breaking verification. The audit chain now commits to a per-entry content hash (entry_hash = SHA-256(seq ‖ prev ‖ content_sha256), bundle format 1.1, back-compatible with 1.0), so redacted content is replaced by its commitment and the chain still verifies. audit_log_hash is invariant under redaction but changes under tampering, so a redaction is an authorized, recorded transformation while a content edit is detected. evidence verify reports which entries are redacted. Encrypted bundles must be decrypted first.

[3.1.0] — 2026-07-18

Additive feature release — no breaking changes. Deepens Boruna’s two moats: verifiable/auditable evidence (standards interop, compliance reporting, observability export, sealed contract/guard verdicts) and agent authoring (exact-signature lookup, a run-and-seal execution cell, an agent corpus). Ideas were mined from adjacent tooling (Temporal/LangGraph/Langfuse/Credo AI/SLSA/ in-toto/Sigstore) and from the agentlanguages.dev peer catalogue, then mapped onto Boruna’s determinism + evidence model.

Added

  • In-toto + DSSE attestation — boruna evidence attest <dir> emits the bundle as an in-toto Statement (predicateType https://boruna.dev/runtime-provenance/v1) wrapped in a DSSE envelope signed with the bundle’s ed25519 key; --verify checks it. Makes runtime-execution provenance consumable by the supply-chain ecosystem (cosign, in-toto-verify). Additive — the native bundle is unchanged. Predicate schema: docs/spec/runtime-provenance-predicate-1.0.md.
  • Compliance-mapping report — boruna evidence report --framework eu-ai-act|nist|iso42001 verifies a bundle, then maps its contents to the specific obligation each helps satisfy (EU AI Act Art. 12/19/26, NIST AI RMF, ISO/IEC 42001), honestly flagging gaps. A technical mapping, not a certificate of compliance.
  • OpenTelemetry export — boruna evidence otel <dir> emits the run as OTLP/JSON spans (no SDK dep, no network) with tamper-evidence attributes (boruna.bundle_hash, audit_log_hash, signature.keyid) and gen_ai.* spans for llm.* calls, so a run surfaces in any OTel backend while linking back to a verifiable record.
  • Sealed contract + guard verdicts — requires/ensures contract checks now record a ContractCheck event (pass and fail) into the hash-chained evidence log. New __builtin_guard(value, passed, label) runs a deterministic output check, traps fail-closed on violation, and seals the verdict — so “the guardrail ran on this model output and returned this verdict” becomes a replayable, tamper-evident fact.
  • std-guard standard library (14th lib) — pure, deterministic output validators (length/range/allow-list/ban-list/refusal-heuristic/json-shape).
  • MCP tools (now 14) — boruna_symbols (exact typed signatures for .ax source) and boruna_run_sealed (compile + run + replay-verify → a verifiable execution record).
  • Quickfix-coverage CI gate — every auto-fixable diagnostic must ship a repair strategy or be explicitly allow-listed.
  • Agent corpus & docs — llms.txt, an .ax teaching primer, a static agent portal manifest, an evidence threat model, and a runtime-execution-provenance positioning doc.

Fixed

  • docs/reference/ax-language.md syntax drift — corrected to the real grammar (records use type, enum variants are unit or single-payload, match arms use bare variant names), verified with boruna lang check.

[3.0.0] — 2026-07-18

Removes the entire HTTP / serving / distributed-execution layer. Boruna is now a local deterministic engine + CLI — no HTTP server. The compiler, VM, orchestrator engine, evidence bundles, deterministic replay, and every local CLI command are unchanged. Breaking, hence the major bump: public CLI commands and a build feature were removed.

Removed

  • The HTTP serving / distributed-execution layer — the coordinator (distributed HTTP server), distributed workers, active-active HA, coordinator mTLS, the workflow dashboard, the evidence web viewer, and the approval console.
  • The serve cargo feature and its server dependencies (axum, hyper, tower, reqwest, rustls, …).
  • CLI commands coordinator, dashboard, worker, and evidence serve; and the --coordinator / --coord-token flags on workflow run/approve/reject/trigger. Approval and trigger gates are still handled locally via boruna workflow approve/reject/trigger + resume.
  • Net: ~11,000 lines removed.

Kept

  • The local engine (boruna-orchestrator, boruna-vm, boruna-compiler), evidence bundles, deterministic replay, capability-policy enforcement, and every local CLI command (run, workflow …, evidence verify/inspect, lang, template, migrate, framework, policy, metrics), plus the MCP server.
  • The http feature — the VM’s outbound net.fetch capability for workflow steps (a workflow capability, not a server).

[2.0.0] — 2026-07-17

First major release. A security-hardening + language-completeness sprint that remediates every finding from a whole-codebase research audit (3 High, several Medium, plus the “statically typed but unchecked” language gaps). It carries deliberate breaking changes — integer overflow and several coordinator/framework defaults now fail closed — hence the major bump. See Breaking changes below; each has a documented migration or override.

Security

  • SSRF hardened in the live HTTP handler. URL safety is now split into a syntactic check and a live DNS-resolution check that rejects every resolved private/loopback IP (IPv4 and IPv6, brackets stripped); redirects are followed through a bounded manual loop that re-validates each hop, so a public host can no longer redirect into the internal network.
  • Coordinator cross-worker claim hijack closed (S6). Completing, failing, or extending a step now requires the caller to own the claim; a non-owner is rejected with 403 coord.claim_not_owned, checked at the trust boundary under the store lock.
  • Coordinator approval-gate forgery closed (S9). An approval gate now mints a per-gate token; approve/reject require it (403 coord.approval_token_invalid).
  • Evidence bundles are now tamper-evident. Verification gained an external anchor (evidence verify --expected-bundle-hash <hex>) plus optional ed25519 manifest signing (--verify-key / --require-signature) and a --require-encryption downgrade guard — a forged but internally-consistent manifest that plain verify accepted is now caught.
  • Content-addressing enforced at the coordinator. A worker’s output_hash is verified against SHA-256(output_json) on completion.
  • XSS fixed in evidence serve — all bundle-derived HTML is escaped.
  • Key material zeroized — evidence DEK/KEK wiped on drop.
  • Path-traversal guards added to template names and storage ref_to_run_id (S3/GCS/Azure).
  • Crafted-SpawnActor DoS fixed — a bad function index fails the actor instead of panicking the VM.

Added

  • User enum construction + real per-variant match tags. Enums could be declared and matched but never constructed; there is now an expression form Enum::Variant / Enum::Variant(payload) (new :: token) that compiles to MakeEnum, and match arms dispatch on the variant’s real declaration index (they previously all collapsed to the first arm).
  • Higher-order / indirect calls via new Op::CallIndirect — a function passed as a value now dispatches correctly (previously hardcoded to fn #0).
  • for loops, Map<K,V> / Fn(..) -> T type annotations, and ensures postconditions.
  • Static arity checking — a direct call to a named function with the wrong argument count is now a compile error.
  • Warn-only static type-consistency checking (E009 warnings). lang check / boruna_check now surface let-annotation and call-argument type mismatches as warnings, without blocking compilation — the first, non-breaking step of a staged rollout toward strict typing.

Breaking changes

  • Integer overflow is now a runtime error (VmError::ArithmeticOverflow), where it previously wrapped (release) or panicked (debug).
  • The coordinator refuses to start on a non-loopback bind without auth. Override for trusted networks: BORUNA_COORD_ALLOW_INSECURE=1.
  • The coordinator rejects an output_hash that doesn’t match output_json (previously trusted).
  • Framework policy defaults fail closed: an empty or malformed policies() now denies (was allow-all). Apps that define no policies() at all keep the allow-all convenience.
  • Codegen rejects element/field/argument counts above 255 with a compile error instead of silently truncating a u8 operand.

Fixed

  • while-body trailing-expression stack leak (one operand leaked per iteration).
  • Documentation/version/count drift (README, CLAUDE.md, stability, stdlib manifest) corrected to the real workspace state.

[1.9.0] — 2026-07-15

Ninth feature minor on the 1.x LTS line. Fourth and final sprint of the agentlanguages.dev competitive-borrow program (Theme D-lite): capability-row inference — the compiler now infers each function’s minimal capability set and flags over-declarations.

Added

  • Capability-row inference + over-declaration check (boruna lang caps). Theme D-lite — borrowed from AILANG (effect-row inference: the compiler computes the minimal capability set and flags over-declaration). New Module::needed_capabilities(func_idx) infers the capabilities a function actually needs — those it invokes directly via CapCall plus everything its transitive callees need (cycle-safe, deterministic). Module::over_declared_capabilities(func_idx) returns the capabilities a function declares (!{...}) but never (transitively) uses — an over-grant of authority (a least-privilege smell; not a correctness bug, since the VM still gates at runtime). The new boruna lang caps <file.ax> [--json] command reports each function’s declared vs. inferred-needed capabilities and exits non-zero when any over-declaration is found, so it can gate least-privilege in CI. Deliberately no information-flow / data-visibility typing (a large type-system addition; deferred). See docs/design-capability-inference.md.

[1.8.0] — 2026-07-15

Eighth feature minor on the 1.x LTS line. Third sprint of the agentlanguages.dev competitive-borrow program (Theme C-lite): the LLM effect now propagates up the call graph and is recorded in evidence.

Added

  • LLM effect propagates up the call graph; model-invoking steps recorded in evidence. Theme C-lite — borrowed from Vera (LLM inference as a tracked typed effect). New Module::transitively_invokes(func_idx, capability) computes whether a function reaches a capability through its call graph (its own declared capabilities, or any function it Calls / SpawnActors, transitively — cycle-safe, order-independent). When a workflow runs, each source-kind step is analysed for transitive llm.call reachability and the sorted list of model-invoking step ids is captured into the evidence bundle as model_invoking_steps.json — checksummed and covered by bundle_hash, so evidence verify fails on tamper. An auditor can now see which steps touched a model, even when the call is indirect through a helper. Deliberately no conformal-prediction / uncertainty quantification (research-grade; deferred). See docs/design-llm-typed-effect.md.

[1.7.0] — 2026-07-15

Seventh feature minor on the 1.x LTS line. Second sprint of the agentlanguages.dev competitive-borrow program (Theme A-lite): runtime-checked contracts with concrete, replayable counterexamples — no SMT.

Added

  • Runtime-checked requires preconditions with counterexample evidence. Sprint 2 / Theme A-lite of the agentlanguages.dev competitive-borrow program — borrowed from Vera/Aver (design-by-contract) and Vow (counterexample = concrete replayable input). A function’s requires <expr> clauses are now compiled to runtime guards checked at entry against the arguments; a violation traps with a new VmError::ContractViolation { message, counterexample } where counterexample is the offending argument list (positional, rendered) — the exact input an auditor needs to reproduce the breach. The failure surfaces with a stable, distinct error_kind contract_violation (retry: no — a violation is a deterministic function of the inputs), and because failed-step errors are recorded in the hash-chained audit log, the counterexample lands in tamper-evident evidence. Reuses the previously-dormant Op::Assert opcode (no bytecode bump); functions without contracts emit no guard. Deliberately NO SMT/Z3 — this stays in Boruna’s concrete-trace + replay philosophy, consistent with the 1.5.0 “Decided” ruling against symbolic model checking. ensures postconditions (needing a result binding at each return) are a documented follow-up. See docs/design-contracts-runtime.md.

[1.6.0] — 2026-07-15

Sixth feature minor on the 1.x LTS line. First sprint of the agentlanguages.dev competitive-borrow program (Theme B): machine-read intent declarations captured into tamper-evident evidence bundles.

Added

  • intent "..." declarations captured into evidence bundles. Borrowed from Pact (intent-in-signature) and Intent/Prove (machine-read purpose), per the agentlanguages.dev competitive research (claudedocs/research_agentlanguages_competitive_2026-07-15.md, Theme B). A function may declare a single machine-read purpose after its signature: fn transfer(x: Int) -> Int !{db.write} intent "Move funds between accounts" { ... }. The clause is optional, order-independent with requires/ensures, and a second intent on one function is a parse error. Intent threads through lexer → AST (FnDef.intent) → codegen → bytecode (Function.intent, additive #[serde(default)] — pre-Sprint-1 modules load with None) and is surfaced in boruna ast --json. When a workflow runs, each source-kind step’s intent is captured into the evidence bundle as intents.json (step_id → declared purpose), covered by the bundle’s checksums and bundle_hash — so evidence verify fails if a captured intent is tampered, and an auditor sees what each step was authorized to do next to what it did. Determinism (§15): intent is replay-verified evidence, already transitively in workflow_hash. See docs/design-intent-evidence.md.

[1.5.0] — 2026-05-20

Fifth feature minor on the 1.x LTS line. Quint-inspired tooling additions: literate workflow specs, ITF trace export, an interactive REPL, random property-based workflow simulation with witnesses, and the __builtin_debug print-and-passthrough helpers. The first bytecode minor bump (1.0 → 1.1) under the additive-opcode contract.

Added

  • Literate workflow specs (boruna literate extract) — borrowed from Quint’s Literate Specifications. A markdown file with <lang> <filename> += code fences (where <lang> ∈ ax|boruna|quint) is the single source of truth for both the audit narrative AND the executable Boruna source. boruna literate extract <file.md> --out-dir <dir> walks the document, validates each fence, and emits per-file outputs that compile and run via the normal boruna run / workflow run paths. Idempotent: re-running produces byte-identical output. Path traversal and absolute paths are rejected at parse time with stable error_kind strings (literate.invalid_fence, literate.path_traversal, literate.absolute_path, literate.invalid_out_dir, literate.io). New module tooling/src/literate/, reusable from any caller including boruna-mcp. Example fixture at examples/literate/hello_literate.md. See docs/design-literate-workflows.md and docs/architecture-literate-workflows.md.
  • ITF (Informal Trace Format) export from evidence bundles — borrowed from Quint / Apalache / the ITF Trace Viewer. boruna evidence inspect <bundle> --itf emits the bundle’s audit log as an ITF v0.15 document on stdout, with one ITF state per audit-log entry and event variant names preserved as #meta.action. --itf is mutually exclusive with --json. Boruna’s internal evidence-bundle format is unchanged — ITF is purely an export. Vendored producer constant ITF_FORMAT_VERSION = "0.15". New module tooling/src/trace/{itf,audit_to_itf}.rs. Spec source: https://apalache-mc.org/docs/adr/015adr-trace.html. See docs/design-itf-traces.md and docs/architecture-itf-traces.md.
  • __builtin_debug(v) / __builtin_debug_msg(msg, v) — print-and-passthrough debug helpers (bytecode 1.1). Borrowed from Quint’s q::debug. The single-arg form prints Value::Display form to stderr and returns the value unchanged; the two-arg form prints <msg> <value>\n. Operational-only — no capability gate, no audit-log event, no replay impact. Implemented as two new opcodes Op::Debug (0xA7) and Op::DebugMsg (0xA8); BYTECODE_VERSION bumped to "1.1". A 1.0 reader presented with either opcode MUST reject with an unknown-opcode error per §1.2(6) of docs/spec/bytecode-1.0.md. See docs/architecture-q-debug.md.
  • boruna repl [file.ax] — interactive REPL for .ax modules. Borrowed from Quint’s quint repl. Loads an optional initial .ax file, evaluates expressions, supports meta-commands :load, :reload, :reset, :type, :env, :help, :quit. Per-input compile + fresh VM (avoids module-frozen-by-construction VM invariants); the synthetic wrapper declares Int return type because the typechecker is currently permissive about return-type unification, and :type reports the post-hoc Value::type_name(). Defaults to --policy deny-all because REPL inputs are non-deterministic. No line-editing dep — uses std::io::BufRead, sufficient for piped agent-driven use. Bytecode 1.1’s __builtin_debug works in the REPL. See docs/architecture-boruna-repl.md.
  • boruna simulate <dir> [--invariant <expr>] [--witnesses name=expr,...] — random property-based workflow simulation. Borrowed from Quint’s quint run. Runs the workflow --max-samples times (1..=100_000, default 1000) and reports invariant violations + per-witness trace frequencies. The invariant / witness DSL accepts status == "...", total_duration_ms < N, step.<id>.status == "...", step.<id>.duration_ms < N, combined via && / || with parentheses. Per project-conventions-2026-04 §15 the simulator’s per-trace WorkflowRunResult is operational-only and never feeds production replay verification. Sequential v1; input fuzzing and parallel execution are documented follow-ups. Stable error_kind strings: simulate.invalid_samples, simulate.invalid_workflow, simulate.invariant_parse, simulate.witness_parse. New module orchestrator/src/simulate/{mod,invariant,witness}.rs. See docs/architecture-boruna-simulate.md and docs/architecture-boruna-witnesses.md.

Decided

  • docs/spec/bytecode-1.0.md v1.1 minor bump — additive opcodes per §1.2(6) of the spec. Updates the version-identifier prose, adds §4.5 “1.1 additions” with the new opcode table entries (Debug 0xA7, DebugMsg 0xA8), and appends a §12 changelog entry. Backwards compatibility within 1.x is preserved: a 1.1 module CAN be rejected by a 1.0 reader (unknown-opcode error), but every 1.0 module continues to load on a 1.1 reader.
  • Apalache-style bounded symbolic model checking is NOT recommended for Boruna. A pull-in of Apalache + Z3 + a Boruna IR → SMT translator would be a multi-engineer-year integration aimed at an audience (consensus-protocol provers) that does not appear in Boruna’s compliance-runtime positioning. Boruna’s concrete-trace + replay + evidence-bundle model is a different design philosophy and stays as-is. See claudedocs/research_quint_borrowable_ideas_2026-05-20.md.

[1.4.0] — 2026-05-17

Fourth feature minor on the 1.x LTS line. Agent-native CLI inspection surfaces, the boruna-lsp language server, and compliance example workflows.

Added

  • Agent-native CLI surfaces — five read-only, --json-capable commands so AI agents can inspect Boruna projects without reading source. Motivated by a competitive review of vercel-labs/zero.
    • boruna lang codes [--json] — emit the registry of stable diagnostic codes (E001–E009) with name, summary, and category. Backed by tooling/src/diagnostics/registry.rs; a drift test keeps the registry 1:1 with the E0NN constants the compiler emits.
    • boruna doctor [--json] — environment and toolchain health: binary version, compiled features, Rust toolchain, data-directory writability, and project-layout detection. Exits 1 if any check fails.
    • boruna workflow graph <dir> [--json] — emit DAG facts for a workflow: nodes (kind, capabilities, dependencies), edges, topological order, roots, and leaves. Exits 1 on a non-DAG.
    • boruna size <file.ax> [--json] — bytecode artifact cost: per-function opcode counts, module-wide totals, and serialized .axbc byte size.
    • boruna skills list / boruna skills get <name> [--json] — embedded, agent-curated documentation (ax-language, cli, workflows, diagnostics) compiled into the binary, usable with no repository checkout.
  • docs/reference/diagnostic-codes.md — human reference for the diagnostic-code registry.
  • boruna-lsp language server — new crate crates/boruna-lsp. A Language Server Protocol implementation for .ax files providing live diagnostics, completion, and formatting in any LSP-capable editor (VS Code, Neovim, …). See docs/guides/lsp.md.
  • Compliance example workflows — three regulated-use-case workflows under examples/compliance/: soc2_audit_workflow (SOC 2 audit trail), hipaa_data_pipeline (PHI redaction + audit log), financial_review_pipeline (dual-control SOX approval gates). See docs/reference/compliance/README.md.

[1.3.0] — 2026-04-30

Stable

  • std-llm is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-llm.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-json is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-json.md; bumps require a 1.x deprecation notice per LTS contract.

Added

  • 27 new language built-in functions — comprehensive string, list, and map operations now available in .ax programs without importing any library:
    • String: __builtin_int_to_string, __builtin_float_to_string, __builtin_string_len, __builtin_string_chars, __builtin_string_contains, __builtin_string_starts_with, __builtin_string_ends_with, __builtin_string_to_upper, __builtin_string_to_lower, __builtin_string_trim, __builtin_string_join, __builtin_string_split, __builtin_string_replace, __builtin_string_slice, __builtin_int_parse, __builtin_float_parse, __builtin_bool_to_string
    • List: __builtin_list_len, __builtin_list_is_empty, __builtin_list_head, __builtin_list_tail, __builtin_list_append, __builtin_list_concat, __builtin_list_reverse
    • Map: __builtin_map_get, __builtin_map_set, __builtin_map_remove, __builtin_map_contains_key, __builtin_map_keys, __builtin_map_values, __builtin_map_len
  • Import resolution — import "std-name" statements in .ax source now resolve at compile time via a source-level preprocessor that inlines the named library from libs/<name>/src/core.ax. No compiler pipeline change required.
  • boruna evidence inspect shows step outputs — for plaintext bundles, evidence inspect <bundle> now reads outputs/<step_id>/result.json and renders a truncated preview (500 chars) per step in text mode; --json mode includes a "step_outputs" key with full parsed content. Encrypted bundles without --decrypt print a hint to stderr.
  • std-json enhancements — json_array(items: List<String>) -> String serializes a list to a JSON array string; int_to_string now calls __builtin_int_to_string (was returning empty string); json_escape now performs proper character-by-character escaping using __builtin_string_chars.
  • std-validation enhancements — string_length now calls __builtin_string_len (was hardcoded 0); added validate_contains, validate_starts_with, validate_ends_with.

1.2.0 — 2026-04-29

Stable

  • std-ui is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-ui.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-validation is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-validation.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-forms is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-forms.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-authz is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-authz.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-http is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-http.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-db is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-db.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-sync is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-sync.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-routing is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-routing.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-storage is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-storage.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-notifications is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-notifications.md; bumps require a 1.x deprecation notice per LTS contract.
  • std-testing is now 1.0-stable. Public surface frozen per docs/reference/stdlib/std-testing.md; bumps require a 1.x deprecation notice per LTS contract.

Added

  • Compliance templates — three pre-built workflow patterns: soc2_audit_workflow (SOC 2 audit trail), hipaa_data_pipeline (PHI redaction + audit log), financial_review_pipeline (dual-control SOX approval gates)
  • Four new example workflows demonstrating stdlib package usage (form_submission_pipeline, data_ingestion_pipeline, api_routing_workflow); closes graduation criterion 1 for all 11 std-* packages
  • docs/reference/stdlib/std-llm.md and docs/reference/stdlib/std-json.md — reference docs closing criterion 4 for std-llm and std-json
  • examples/workflows/llm_content_generator/ and examples/workflows/json_data_transformer/ — example workflows closing criterion 1 for std-llm and std-json
  • boruna evidence diff — compare two evidence bundles side-by-side. Reports differences in step outputs, audit event counts, workflow metadata, and verification status. --json flag for machine-readable output.
  • boruna workflow eval — run the same workflow against two LLM provider configs and compare evidence bundles; reports per-step output agreement and timing

Changed

  • Improved error messages: boruna lang check now suggests the nearest variable name for E003 errors and the nearest function name for E004 errors using edit-distance-1 matching; type-conversion hints for common E009 mismatches (Int↔String, Bool↔Int); E001 lexer errors now include a source pointer line; E002 parse errors append a common-cause hint; E007 capability violation message now names the offending capability and action.
  • Better lang repair: repair now handles E003 near-miss rename patches via the tooling suggestion pipeline; new RepairStrategy::Conservative applies only High-confidence patches, skipping Medium/Low (safe for CI auto-repair); bottom-up patch ordering was already in place and verified correct.

1.1.0 — 2026-04-29

Added

  • Capability call markers in MCP progress notifications (post1-T-2.2). boruna_run streaming progress events now carry a message field when a capability call fires during an execution slice: "cap: llm.call" for a single call or "caps: llm.call, net.fetch" for multiple. Slices with no capability calls continue to omit message (no noise for pure compute). MCP clients that display live execution status can now surface "calling llm.call…" feedback without polling. Backward-compatible: existing clients that ignore message see no change.

  • Web evidence bundle inspector (post1-T-4.4). New boruna evidence serve <bundle-dir> [--port <port>] subcommand (requires serve feature) starts a local axum HTTP server on port 4444 and opens the browser automatically. Pages: /bundle (overview + verification status + file checksums), /audit (hash-chained event timeline), /outputs (per-step result JSON accordion), /api/bundle (raw JSON dump). Bundle data is loaded once at startup; verification runs via the existing verify_bundle() path and surfaces PASS/FAIL inline. Works offline — no external CDN dependencies.

  • BYOH reference handler library (post1-T-1.2). Four new reference CapabilityHandler implementations under examples/llm_handlers/ joining the existing OpenAI example: Anthropic (Messages API), Ollama (local-LLM, deterministic-with-seed), vLLM-and-OpenAI-compatible (one handler covers vLLM/OpenRouter/ Together/Groq/LiteLLM), and AWS Bedrock (skeleton using the AWS SDK because hand-rolling SigV4 adds nothing illustrative). Each is a self-contained ~80–120-LOC copy-and-tweak template with a README documenting auth, response shape, determinism options, and what the reference deliberately omits (multi-provider routing, streaming, cost accounting, etc.). New umbrella examples/llm_handlers/README.md indexes the library and cross-references the built-in LlmRouterHandler (sprint 0.4-S13). New providers.toml.example documents a config-schema convention integrators can adopt for declarative router setup (Boruna does not parse this file); new router_setup.rs shows a reference parser that turns the toml into an LlmRouterHandler.

    This expansion is faithful to the BYOH design contract shipped in 0.3-S8: Boruna does not ship default handlers in core. Each reference is integrator-copyable code, not a Cargo dep. Auditing the original T-1.2 plan (“ship a boruna-effect-providers adapter crate”) against the shipped BYOH guide flagged the premise conflict before any code was written; the reframed scope delivers the spirit of “more provider on-ramps” without violating the contract.

Changed

  • BundleStorage trait promoted to public 1.x API surface. BundleStorage, StorageRef, StorageError, LocalFs, and the from_uri dispatcher in boruna_orchestrator::audit::storage shipped behind #[doc(hidden)] while the shape was still being validated against remote impls. With T-3.1 (S3), T-3.2 (GCS), and T-3.3 (Azure Blob) all landed and exercising the trait identically, the shape is stable and the hidden attribute is removed. The per-adapter modules (storage_s3 / storage_gcs / storage_azure) ship without #[doc(hidden)] from the start, so this change is purely a rustdoc visibility tweak — no API breakage. StorageError is now #[non_exhaustive] so future variants are additive. Backend kind strings (s3.transient, azure.permanent, etc.) are also additive — integrators switching on kind should treat unknown values as transient (retryable). New top-level concept page at docs/concepts/bundle-storage.md covers the shared contract; the per-provider operator guides remain in docs/guides/bundle-storage-{s3,gcs,azure}.md.

Decided

  • Stdlib graduation tracker (post1-T-3.4). Assessed all 11 std-* packages against the 4-criterion graduation checklist. Zero packages graduate to 1.0 this cycle. Two criteria fail uniformly: none of the packages is referenced from any examples/workflows/*, and none has a docs/reference/stdlib/<name>.md reference page. Per-package decisions and per-criterion notes are recorded in docs/stdlib-graduation-tracker.md. Closing the gates is filed as Wave-3 follow-up work.

Added

  • Azure Blob Storage adapter for BundleStorage (post1-T-3.3, Wave 3). The --bundle-storage azblob://account/container[/prefix] URI now constructs an Azure Blob Storage adapter when the binary is built with the azure feature (cargo build --features boruna-cli/azure). Same shape as the T-3.1 / T-3.2 adapters, also backed by object_store (with the azure feature toggled). URI shape encodes both the storage account and the blob container so an operator can grep their config and see exactly which account a bundle landed in. Auth via standard AZURE_STORAGE_* env vars (account key, SAS, or service-principal OAuth); AzureBlobBucketBuilder::with_use_emulator(true) switches the SDK into Azurite-emulator mode for local testing. Off by default. When the azure feature is OFF, azblob:// URIs reject at parse time with the actionable-message pattern S3 and GCS use. Backend errors surface with stable error_kind strings (azure.transient, azure.permanent, azure.runtime, azure.unexpected_key). 18 unit tests cover URI parsing, object-path concatenation, ref-to-run-id extraction, and error classification. An Azurite-backed integration test is deferred — Azurite requires SharedKey-signed container creation and object_store doesn’t expose a create_container primitive; pulling in the full azure-storage crate or implementing SharedKey signing for one test wasn’t a proportionate cost. See docs/guides/bundle-storage-azure.md. All three remote schemes (S3, GCS, Azure) now ship — the BundleStorage trait can graduate from #[doc(hidden)] to pub in a follow-up.
  • GCS adapter for BundleStorage (post1-T-3.2, Wave 3). The --bundle-storage gs://bucket[/prefix] URI now constructs a Google Cloud Storage adapter when the binary is built with the gcs feature (cargo build --features boruna-cli/gcs). Same shape as the T-3.1 S3 adapter, also backed by object_store (with the gcp feature toggled). Auth via standard GOOGLE_SERVICE_ACCOUNT / GOOGLE_APPLICATION_CREDENTIALS env vars; GcsBucketBuilder::with_endpoint lets integration tests point at fake-gcs-server. Off by default. When the gcs feature is OFF, gs:// URIs reject at parse time with the same actionable-message pattern S3 uses. Backend errors surface with stable error_kind strings (gcs.transient, gcs.permanent, gcs.runtime, gcs.unexpected_key). Integration tests behind the gcs-it feature spin up fsouza/fake-gcs-server via a custom testcontainers Image (testcontainers-modules has no GCS module) and self-skip when Docker is unreachable. See docs/guides/bundle-storage-gcs.md.
  • S3 adapter for BundleStorage (post1-T-3.1, Wave 3). The --bundle-storage s3://bucket[/prefix] URI now constructs a real remote-storage adapter when the binary is built with the s3 feature (cargo build --features boruna-cli/s3). Backed by the Apache Arrow object_store crate’s aws feature — works against AWS S3, MinIO, Cloudflare R2, Backblaze B2, and LocalStack via the standard AWS_* environment variables (including AWS_ENDPOINT_URL for non-AWS endpoints). The adapter bridges the sync BundleStorage trait against the async SDK with a per-instance current-thread tokio runtime; bundle reads materialize into a local cache directory rooted at BORUNA_BUNDLE_CACHE (defaults to <temp>/boruna-bundle-cache). When the s3 feature is OFF, s3:// URIs reject at parse time with an actionable message that points operators at the feature flag — never silently ignored. gs:// (T-3.2) and azblob:// (T-3.3) remain reserved for upcoming adapters. Backend errors surface with stable error_kind strings (s3.transient, s3.permanent, s3.runtime, s3.unexpected_key). MinIO-backed integration tests live behind the s3-it feature and self-skip when Docker is unreachable. See docs/guides/bundle-storage-s3.md.
  • boruna evidence rotate-kek (post1-T-2.4) re-wraps the DEK of one or more encrypted evidence bundles under a new KEK. Operations are manifest-only — per-file ciphertext stays valid because the DEK itself is unchanged. Supports single-bundle and batch (directory) modes; batch mode runs in parallel via rayon, bounded by --parallelism N (default min(8, num_cpus)). --dry-run validates without writing. --kek-id-from <id> defends against accidental double-rotation in mixed-state batches. New Envelope::rewrap API on the encryption module exposes the same primitive to library consumers. See docs/guides/kek-rotation.md.
  • Pluggable evidence-bundle storage trait BundleStorage and a LocalFs adapter (post1-T-2.3). boruna workflow run --record now accepts --bundle-storage <uri> (or BORUNA_BUNDLE_STORAGE env var); when set, the finalized bundle is copied to the configured backend after the local write succeeds. Storage failure is logged but never fails the workflow — the local bundle remains the authoritative record. Only the local:<root> scheme ships in this release; s3://, gs://, azblob:// are reserved for Wave 3 adapters and reject at parse time. The trait is #[doc(hidden)] until at least one remote adapter ships.
  • boruna_run MCP progress notifications are now part of the 1.x LTS-stable surface (post1-T-1.1). When a client supplies the standard MCP progressToken in the request’s _meta field, the server drives the VM in ~100k-opcode slices and emits notifications/progress events between slices with the cumulative step count. The underlying mechanism shipped in sprint 0.4-S6; this entry formalizes the wire contract and adds reference docs in docs/reference/mcp-server.md § “Progress notifications”.
  • Worker capability advertisements now carry an optional version (post1-T-1.3). RegisterRequest.advertised_capabilities accepts either a bare string (legacy worker, normalized by the coord to the coord’s current Capability::version()) or an explicit {name, version} object. Steps requiring a capability whose version no registered worker advertises now surface coord.capability_version_mismatch (HTTP 409) on the claim response, rather than silently long-polling. The W3-A silent-skip path still applies when a worker is missing the capability NAME entirely. Documented in docs/reference/error-kinds.md and noted in docs/spec/workflow-dag-1.0.md.

Changed

  • Parser and typeck error messages now include did you mean: 'kw'? suggestions for typos within Levenshtein distance 1 of a known keyword (parser) or in-scope identifier (typeck). Suggestions are appended only when a single unique candidate exists, so noisy ambiguous suggestions are intentionally suppressed. The single-line prefix (undefined variable: foo) remains unchanged so existing diagnostic-collector parsers continue to classify errors correctly.

Added

  • docs/post-1.0/README.md describing how post-1.0 work is tracked on GitHub: branch naming, labels, and the project-board scheme.
  • docs/branch-policy.md documenting the master (1.x LTS) vs. 0.7.x (speculative) branch topology and cross-merge rules.
  • GitHub labels wave-1…wave-4, branch-master, branch-0.7.x, post1-execution for filtering post-1.0 PRs and issues.
  • Bench compare CI job and .github/scripts/bench_compare.py — PR-time perf-regression detection that runs the criterion harness on PR base and head, posts a sticky comment with per-benchmark deltas, and fails on ≥10% mean regression. Non-blocking by default (not in the required-status-checks set). See CONTRIBUTING.md § “Reading the bench-compare PR comment”.
  • Smoke test (musl) CI workflow and .github/scripts/smoke_musl.sh — automated container-based smoke tests for the x86_64-unknown-linux-musl and aarch64-unknown-linux-musl release artifacts. Runs on every v*-rc* tag (or via workflow_dispatch for an existing tag), verifies SHA-256, launches the binary under alpine:3.19, runs the llm_code_review example end-to-end, verifies the evidence bundle, and opens a PR with docs/release-smoke-tests/<tag>-musl-<arch>.md reports. The aarch64 leg runs under qemu-user-static and is explicitly NOT a real-hardware smoke; that remains operator-side.
  • boruna run --watch — re-execute a .ax file on every change. Debounces filesystem events to 200ms, prints a ── reloading <path> at HH:MM:SS ── separator before each rerun, and tolerates per-run errors so the watcher keeps running across fix-and-save cycles. See docs/reference/cli.md § Watch mode.

1.0.0 - 2026-04-28

First stable release. The 1.x LTS contract takes effect from this tag forward — every 1.0 .ax program, workflow.json, evidence bundle, MCP integration, and CLI invocation is committed to keep working on every 1.y release per docs/lts.md §B.

Same surface as 1.0.0-rc3. No code changes between rc3 and this GA cut; this tag exists to crystallize the 1.0 LTS commitment and ship final-named binaries.

The four formal versioned specifications frozen at 1.0:

For the full feature scope shipped between v0.5.0 and v1.0.0, see the [1.0.0-rc1], [1.0.0-rc2], and [1.0.0-rc3] sections below.

Decided

  • 1.x LTS contract is now in force. docs/lts.md §B surfaces are stable through 2027-11 (active) / 2028-05 (security). Surfaces classified Experimental in docs/stability.md remain Experimental within 1.x; pin to a specific Boruna release tag if your integration depends on those.

[1.0.0-rc3] - 2026-04-28

Theme: final GA-readiness polish. rc2 shipped W6 (mTLS + bundle encryption) and W7 (security-review closures). rc3 folds the W8-W11 GA-polish work into a tagged candidate so operators have a single artifact representing the actual GA candidate to soak. Highlights:

  • 4th formal versioned specification (bytecode 1.0) publishes alongside the existing three (.ax language, workflow DAG, evidence bundle), all locked behind reader constants per docs/lts.md §B.
  • Algorithm-gate enforcement in evidence bundle decryption (W7 NEW-1): Envelope::unwrap now rejects bundles declaring algorithm ≠ aes-256-gcm with evidence.unsupported_algorithm, before any KEK-related work — matches the spec’s reader contract.
  • CI hardening: bench harness compiles on every PR (W8); examples run end-to-end + verify on every PR (W9-D); parallel-test flakes fixed (W10).
  • Operator-facing GA-cut tooling: scripts/pre-release-check.sh is the single command that confirms GA-readiness before tagging (W11-A).
  • CHANGELOG-driven release notes (W9-B): the GitHub Release page body is now the CHANGELOG section for the tag, not auto-generated commit noise. First release using this flow is rc3 itself.

After rc3 soak, the v1.0.0 GA tag is a 5-min coding step: bash scripts/pre-release-check.sh 1.0.0 → bump to 1.0.0 → tag → push.

Added

  • Versioned bytecode 1.0 specification at docs/spec/bytecode-1.0.md (sprint W9-A). bytecode_version: "1.0" exposed via boruna_bytecode::BYTECODE_VERSION. Locks the on-disk module format, opcode table, value model, capability table, and determinism contract for the 1.x line. Forward-compat: 1.x VMs accept any 1.y bytecode module.
  • CHANGELOG-driven GitHub Release notes (sprint W9-B). The release pipeline now extracts the CHANGELOG section for the current tag and uses it as the GitHub Release body instead of auto-generating from commits. Operators MUST update CHANGELOG.md before tagging — empty section fails the release loudly. Improves release-page readability for integrators.
  • End-to-end smoke gate for example workflows in CI (sprint W9-D). Each example workflow under examples/workflows/ now runs to completion with --policy allow-all --record and the produced bundle is evidence verify-ed on every push/PR. Catches integration regressions where DAG validation passes but execution fails.
  • cargo bench --no-run gate in CI (sprint W8). The criterion bench harness now compiles on every push/PR so refactors that break bench compilation surface at PR time instead of at the next operator-run baseline.
  • Pre-release validation script at scripts/pre-release-check.sh (sprint W11-A). Read-only script the operator runs before tagging that confirms repo state, version alignment, CHANGELOG coverage, all spec constants, every CI gate, and the examples smoke flow.
  • evidence.unsupported_algorithm typed error (sprint W7 NEW-1). Envelope::unwrap now rejects bundles with an algorithm field other than aes-256-gcm BEFORE any KEK work — matches the evidence-bundle-1.0.md reader contract. Closes the spec/code gap flagged by the W7 follow-up security review.
  • Smoke-test report for v1.0.0-rc2 macOS arm64 artifact at docs/release-smoke-tests/v1.0.0-rc2.md (sprint W9-C). End-to-end verification of the published GitHub Releases binary; pre-GA sign-off for the macOS arm64 target. Linux musl targets remain operator smoke tests on real hardware.
  • 9 missing MCP-layer error_kind strings added to docs/reference/error-kinds.md (sprint W7 NEW-2): closes the taxonomy completeness gap flagged by the W7 follow-up security review. The doc now enumerates 36+ stable error_kind strings across coord.*, evidence.*, workflow.*, policy.*, and MCP-layer namespaces.

Changed

  • scripts/ci.sh refreshed (sprint W11-A) to match the current .github/workflows/ci.yml: clippy --all-targets (W1-A), serve-feature clippy run, bench compile gate (W8).
  • docs/INTEGRATION_GUIDE.md v0.1.0 references replaced with v1.0-GA-aware framing (sprint W10 H-1). The body of the guide remains structurally accurate for v1.0; only the trailing “What Boruna Does Not Do” section was patched.
  • docs/FRAMEWORK_API.md version label dropped (sprint W10 H-2). The framework crate is in the Experimental stability tier per docs/stability.md; the doc now cross- links to that tier definition + the LTS contract instead of carrying a misleading (v0.1.0) heading next to the workspace’s 1.0.0-rc tag.

Decided

  • Cut a third release candidate (v1.0.0-rc3) instead of GA directly (sprint W11). rc2 was published before W7-W11 work landed; cutting GA on current master would skip the soak window entirely and lock the LTS contract on unverified-in-field surfaces (notably the W7 NEW-1 algorithm gate change in Envelope::unwrap). rc3 represents the actual GA candidate; soak runs against rc3, then GA cut.

[1.0.0-rc2] - 2026-04-28

Theme: GA polish. rc1 shipped the post-v0.5 sprint cycle (W1-W6: versioned spec freezes, coord HA, mTLS, capability tagging, blob GC, scaffold, perf baselines, LTS commitment, migration tooling, bundle encryption). rc2 closes the security review remainder before the v1.0 GA tag: explicit GA decision on TLS 1.2 (kept, with rationale); spec amendment documenting the optional encryption block in evidence bundles; canonical error_kind taxonomy reference; mTLS guide updated with revocation + non-ASCII CN limitations; evidence inspect plaintext-leak gate test; algorithm-gate enforcement in Envelope::unwrap matching the spec’s reader contract; stdlib version policy clarified.

Added

  • Canonical error_kind taxonomy reference at docs/reference/error-kinds.md (sprint W7). All stable error_kind strings enumerated with HTTP status, sprint origin, and caller-facing meaning per LTS §B.6. Cross-linked from the policy-schema reference and the evidence bundle spec; integrators may switch on these strings.
  • Evidence bundle encryption block documented in spec (sprint W7, finding M-1). docs/spec/evidence-bundle-1.0.md now formally describes the optional encryption field added in W6-B: field shape, AES-256-GCM algorithm pin, per-file nonce derivation, replay-verified vs. operational classification, and the reader contract (1.x readers WITH KEK, 1.0 readers without). Additive to format_version: 1.0; no version bump.
  • mTLS limitations documented (sprint W7, findings M-4 / M-5). docs/guides/coord-mtls.md gains a “Limitations” section calling out the absence of CRL/OCSP revocation and recommending short-lived (≤24h) certs as the v1 mitigation, plus a “CN comparison semantics” subsection documenting the ASCII-only eq_ignore_ascii_case fold (non-ASCII CNs are case-sensitive; no Unicode normalization).
  • Mutual TLS auth + per-worker client certificates (sprint W6-A). Operators can now require X.509 client certs on the coord HTTP surface via --tls-cert, --tls-key, and --tls-client-ca. Workers present client certs via --tls-cert / --tls-key / --tls-server-ca. The cert subject CN drives worker identity; mismatch with a body worker_id returns coord.identity_mismatch. mTLS is additive: shared-secret bearer auth (sprint 0.5-S3) continues to work unchanged. Operator guide: docs/guides/coord-mtls.md. New error_kind: coord.identity_mismatch.
  • Evidence bundle encryption (sprint W6-B). Operators can now opt into AES-256-GCM envelope encryption for evidence bundles via boruna workflow run --record --encrypt-bundle with the KEK supplied via --bundle-encryption-key <hex> or BORUNA_BUNDLE_KEK env. Per-bundle data keys (DEK) are wrapped with the KEK; bundle.json carries the wrapped DEK and algorithm metadata. verify_bundle auto-detects encryption and decrypts before integrity check. Backwards- compat: unencrypted bundles continue to work. KEK lifecycle is the operator’s responsibility — Boruna does not manage keys. New error_kinds: evidence.encryption_key_required, evidence.encryption_key_mismatch, evidence.cipher_tag_invalid. Threat model: docs/design-bundle-encryption.md.

Decided

  • TLS 1.2 remains enabled in W6-A mTLS (sprint W7). Decision: the default rustls 0.23 + aws_lc_rs configuration restricts TLS 1.2 to AEAD-only ciphers (no CBC, no RC4, no export-grade), which is considered safe for 1.0 GA. Operators wanting TLS-1.3-only can build with a forked rustls feature set; the default ships TLS 1.2 for compatibility with older HTTP load balancers and worker hosts. Rationale: the cryptographic surface (AEAD ciphers, ECDHE key exchange) is the same as TLS 1.3 for the practical attack model; forcing TLS-1.3-only would block deployment on systems with older client/proxy stacks. Re-evaluate at 2.0 if TLS 1.3 adoption is universal.

[1.0.0-rc1] - 2026-04-28

Theme: 1.0 release candidate. This is the first 1.0 release candidate. Surfaces listed in docs/lts.md section B are now LTS-protected under the long-term-support contract that takes effect at 1.0 GA. Three formal versioned specifications are published and frozen at 1.0: the .ax language, the workflow DAG schema, and the evidence bundle format. The distributed-execution stack from v0.5.0 ships HA-ready (multi-coord active-active behind a load balancer or via worker URL failover). Workers can advertise capability subsets so heterogeneous fleets are supported. Operators get the boruna new interactive scaffold, boruna migrate for upgrading legacy artifacts, and boruna evidence gc-blobs for blob storage cleanup. Performance baselines are published with 1.x budget commitments.

Added

  • boruna migrate subcommand (beta) (sprint W5-C). Migrators for evidence bundles (synthesize missing bundle.json for legacy v0.5.0-and-earlier bundles) and workflow.json (add schema_version: 1 when missing). --dry-run previews; --in-place modifies the input

  • boruna migrate subcommand (beta) (sprint W5-C). Migrators for evidence bundles (synthesize missing bundle.json for legacy v0.5.0-and-earlier bundles) and workflow.json (add schema_version: 1 when missing). --dry-run previews; --in-place modifies the input directly; default writes a .migrated sibling. Beta status: the migrator coverage will expand in 1.x as breaking changes accumulate. Operator guide: docs/guides/migration.md.

  • Performance benchmarks baseline (sprint W5-A). New benches/ workspace member with criterion-based benchmarks for compile time, VM throughput, and evidence bundle write/verify. Documented baseline + 1.x performance budget commitments at docs/PERFORMANCE.md. Benches are not gated in CI; run locally via cargo bench -p boruna-benches.

  • Long-term-support contract for 1.x (sprint W5-B). New docs/lts.md documents the support windows (1.x active for 18 months from 1.0 GA, security-supported for 24 months; 0.x EOL on 1.0 GA), the LTS-protected surface (.ax language_version: "1.x", workflow DAG schema, evidence bundle format, MCP protocol_version: 1 responses, CLI commands and flags, error_kind strings, HTTP API wire format), the deprecation policy (announce in 1.y → runtime warning → 6-month notice → migration tooling) for breaking changes in 2.x, the security-fix backport policy (CVSS v4, CRITICAL/HIGH within 7 days), and the 12-month end-of-life procedure. docs/stability.md cross-links to the LTS contract and clarifies which tiers are LTS-protected. The README gains an LTS line near the badges. SECURITY.md gains a backport-policy section. Doc-only sprint, no code changes.

  • Worker capability tagging (sprint W3-A). Workers may advertise a SUBSET of the coord’s capability set via --advertise-caps net.fetch,db.query; coord routes only steps whose policy-required capabilities are a subset of the worker’s advertised set. Backwards-compatible: workers omitting the flag behave as before (full fleet). New error_kind: "coord.unknown_capability" rejects registration with unknown capability names. Operational metadata only — placement filter, not a security gate; the VM’s capability gateway remains the authority.

  • Blob GC (sprint W3-B). New boruna evidence gc-blobs command sweeps orphan content-addressed blobs from the data-dir’s blobs/ tree (output blobs no longer referenced by any step checkpoint). --dry-run reports without deleting; --json emits a structured report. Closes the 0.5-S7 accepted limitation around manual blob cleanup. Library APIs BlobStore::find_orphans, BlobStore::delete, and RunCheckpointStore::all_referenced_blob_hashes are also exposed for future coord-side periodic-sweep wiring.

  • boruna new interactive scaffold (sprint W3-C). Wraps the existing template engine with stdin-driven prompting. Walks the user through template selection, target dir, and per-template variables; confirms before writing. --no-input mode is CI-safe (errors on missing defaults rather than silently filling). Refuses to overwrite non-empty target dirs without --force.

  • Coordinator HA / failover (sprint W2). Multiple boruna coordinator serve processes can run against the same SQLite data-dir for active-active HA. Workers accept comma-separated URLs in --coordinator and try them in order at registration time, sticking to the first reachable one. New GET /api/health endpoint returns {status, boruna_version, capability_set_hash, uptime_ms} and bypasses bearer auth so external load balancers can probe without holding the secret. Deployment topologies and failure-mode walkthroughs are documented at docs/guides/coord-ha.md.

  • Versioned workflow DAG schema (sprint W4). New schema_version: 1 field required on every workflow.json. Spec at docs/spec/workflow-dag-1.0.md. boruna_orchestrator::WORKFLOW_DAG_SCHEMA_VERSION = 1 exposed for compatible readers. Forward-compat: 1.x readers accept any 1.y workflow (additive fields ignored).

  • CI clippy gate now uses --all-targets (sprint W1-A). All three clippy invocations in .github/workflows/ci.yml now include --all-targets so test-code lint regressions surface at PR time instead of only at release-runner time. Filed as a followup in the B-2 retro after the workspace --all-targets sweep landed.

  • Formal versioned .ax language specification at docs/spec/ax-language-1.0.md (sprint W1-B). language_version: "1.0" exposed via boruna_compiler::LANGUAGE_VERSION. New docs/spec/README.md indexes versioned specs. The narrative reference (docs/reference/ax-language.md) cross-links the spec.

  • Versioned evidence bundle format with format_version: "1.0" in bundle.json (sprint W1-C). Forward-compat reader gate rejects bundles from incompatible major versions; same-major bundles are accepted with unknown fields ignored. Spec: docs/spec/evidence-bundle-1.0.md.

Changed

  • BREAKING: Evidence bundles now require a top-level bundle.json manifest (sprint W1-C). Legacy bundles from v0.5.0 and earlier must be migrated (migration tool planned for sprint W5-C; until then, re-record against a current binary).
  • BREAKING: workflow.json files without schema_version are now rejected (sprint W4). All bundled examples updated. Operator action: add "schema_version": 1 to existing workflow definitions before upgrading.

Decided

  • 1.x is the long-term-support line. At 1.0 GA, the surfaces listed in docs/lts.md section B are LTS-protected for the full 1.x line: every 1.0 .ax program, workflow.json, evidence bundle, MCP integration, and CLI invocation continues to work unchanged on every 1.y. Active support runs 18 months from 1.0 GA, security support 24 months. Breaking changes ride the 2.0 boat with at least 6 months of deprecation notice and migration tooling for any mechanically-derivable upgrade. Internal Rust APIs, default values, and logging output formats are explicitly out of scope — Boruna ships a CLI + binary, not a Rust library. See docs/lts.md for the full contract.

0.5.0 - 2026-04-28

Theme: distributed execution. Boruna can now run a fleet of worker processes coordinated by a single HTTP coordinator, drive workflows over the wire from CI runners that don’t share a data-dir, handle large LLM step outputs without bloating the SQLite store, and serve human-in-the-loop and webhook-driven gates against a remote cluster. Read paths are consistent across in-process resume, evidence-bundle creation, dashboard rendering, and the step_input builtin — every persistence reader of step outputs goes through the same accessor.

The 0.5-S2a → 0.5-S2f sub-sprint cycle landed during 0.4.x and is included in this tag for the first time as a versioned release (the distributed-execution stack: claim/lease persistence, coordinator/worker HTTP MVP, lease-expiry sweep, coord+dashboard listener-merge, workflow run --submit-only, coordinator wait).

Added

  • Workspace clippy --all-targets is clean (sprint B-2). Pre-existing test-code lints from rustc 1.91+ in 4 crates (boruna-bytecode, boruna-vm, boruna-framework, boruna-compiler) plus a few in production paths cleared in one sweep. Auto-fix handled needless_borrows_for_generic_args, manual_contains, clone_on_copy, for_kv_map, manual_is_multiple_of. Manual fixes for module_inception (4× tests.rs files, #[allow] on inner mod), type_complexity (3 sites in llmvm/capability_gateway tests, factored to a RecordedCalls type alias), approx_constant (test fixture used 3.14 for arbitrary roundtrip — replaced with 2.5), await_holding_lock (existing test had a MutexGuard whose binding scope spanned an .await across an explicit drop(); rebound inside a block scope so drop is automatic before any await).

  • Dashboard renders step outputs with blob-aware fallback (sprint 0.5-S7b). The per-run detail HTML page gains an Output column. Inline outputs render in a <code> block truncated to 256 chars; blob-stored outputs render [blob: <hash[..16]>…] linked to the S7 /api/runs/{run_id}/blobs/{hash} route, without slurping the bytes into the dashboard render. Pending/Running/paused steps show —. Reads route through RunCheckpointStore::read_step_output for inline cases (the same accessor used by the resume and evidence-bundle paths). The JSON detail endpoint (GET /api/runs/{id}) is unchanged — StepCheckpoint already serializes both output_json and output_blob_ref fields, so programmatic consumers can branch on the shape directly. 3 new HTML rendering tests. See docs/design-dashboard-blob-render.md.

  • Output blob references for large step outputs (sprint 0.5-S7). Step outputs whose JSON encoding exceeds 64 KiB are now offloaded to a content-addressed blob store at <data-dir>/blobs/<aa>/<hash>, keyed by SHA-256. The step_checkpoints.output_blob_ref column carries the hash; the inline output_json column is left NULL when the blob path is used. Mutually exclusive: at most one of the two columns is populated for any terminal-state row. Audit hashes are unchanged — the ref IS the existing output_hash, so evidence-bundle replay across pre-S7 and post-S7 runs produces byte-identical hash chains. New schema migration v3 → v4 (additive ALTER TABLE ADD COLUMN, no table rewrite). New coordinator HTTP route GET /api/runs/{run_id}/blobs/{hash}, bearer-gated and run-scoped (the route only serves bytes if the requested hash is referenced by a checkpoint under the given run_id, preventing the route from acting as a generic blob server). New error_kind taxonomy: coord.blobs.bad_hash (400) and coord.blobs.not_found (404). Threshold is hard-coded for the sprint (no Policy knob); a future sprint may make it configurable. 27 new unit tests (15 blob_store, 12 persistence)

    • 5 new coord handler tests. See docs/design-output-blob-refs.md and docs/architecture-output-blob-refs.md.
  • Distributed approval-gate / external-trigger (sprint 0.5-S6). Two new operator-facing routes — POST /api/runs/{run_id}/approve and POST /api/runs/{run_id}/trigger — bearer-gated by the same auth middleware as worker endpoints. CLI flags --coordinator <url> + --coord-token added to boruna workflow approve|reject|trigger so CI runners can drive remote runs without shared data-dir. The wait driver (advance_run_one_tick) now opens approval / trigger gates when their dependencies complete, and closes them when the decision sentinel arrives in metadata.approvals / metadata.triggers — same synthesized output shape as the in-process resume sentinel pass so a run approved via either route hashes to the same evidence bundle. Five handler unit tests + three advance-loop unit tests + one end-to-end CLI integration test. New error_kind taxonomy entries: coord.approve.invalid_state, coord.approve.bad_payload, coord.trigger.invalid_state, coord.trigger.bad_token, coord.trigger.bad_payload. See docs/design-0.5-s6-distributed-approval-trigger.md.

  • boruna workflow run --coordinator <url> (sprint 0.5-S4). Submits a workflow over HTTP to a remote coordinator and polls for terminal status — eliminates the shared-data-dir requirement for CI workflows. Workflow definition + every Source-kind step’s .ax body are inlined into the submit payload. Bearer token via --coord-token or the BORUNA_TOKEN env var. Exit codes match coordinator wait: 0 Completed, 1 Failed, 2 timeout / submit-failed. Two new coordinator HTTP routes: POST /api/runs/submit and GET /api/runs/{run_id}/status, both bearer-gated by the same auth middleware as worker endpoints. Status reads fold advance_run_one_tick into the request so the operator’s poll IS the wait driver. Six handler unit tests + three end-to-end CLI integration tests. New error_kind taxonomy entries: coord.submit.invalid_workflow, coord.submit.bad_payload, coord.runs.not_found. See docs/design-0.5-s4-coordinator-flag.md.

  • Sprint A debt cleanup (preceding commit chore/0.5-debt-cleanup-2). Five carried-forward debts cleared in one pass: eliminate unsafe { env::set_var } (production CLI + 2 tests) by threading env name explicitly through resolve_data_dir/metrics::export; new error_class::TRANSIENT_NETWORK taxonomy entry detected from both VmError::AssertionFailed and wire-level error_msg strings; AuditLog::from_entries_verified called at evidence-bundle creation to catch direct sqlite3 tamper of metadata.audit_log; new Prometheus boruna_workflow_run_duration_seconds histogram for p50/p95/p99 dashboards; five drift-detection tests for docs/reference/policy.schema.json (caught real drift — schema was missing step.input capability, fixed).

  • Carried-debt cleanup pass (preceding session). Three small fixes from earlier-sprint adversarial-review findings that hadn’t been addressed:

    • Audit chain wait-driven terminating event. New WorkflowRunner::append_wait_terminal_audit_event emits WorkflowCompleted to the audit chain when the wait driver reaches Completed or Failed terminal status. Idempotent — re-invoked waits don’t double-emit. Closes the gap from 0.5-S2f where submit-only emitted WorkflowStarted with no terminating entry. 3 new unit tests.

    • Two-concurrent-waits integration test. CORR-6 from 0.5-S2f adversarial review. Locks the design intent: two coordinator wait processes against the same run_id both converge to exit 0 because the underlying insert_pending_step_if_absent and requeue_failed_step_for_retry primitives are INSERT … ON CONFLICT DO NOTHING (race-safe).

    • Submit-only --concurrency warning. Adversarial finding F3 from 0.5-S2e. boruna workflow run --submit-only silently ignored --concurrency because parallelism in distributed mode is controlled by the worker pool, not the in-process wave loop. Now emits a clear stderr warning at submit time so operators know.

  • Path-resolution failure-mode prevention (this session). After the multi-sprint parallel-agent attempt that wrote files to the wrong worktree (because absolute paths in agent prompts bypass the cwd redirection of isolation: "worktree"), this session hardens the workflow:

    • New docs/AGENT-PROMPT-TEMPLATE.md — reusable skeleton for parallel worktree-agent prompts. Bakes in the relative-path discipline + a worktree-verification block that agents must run before any file edit.
    • New CLAUDE.md “Parallel-Agent Best Practices” section documenting the failure mode and the required prevention.
    • New project convention #31 — “Parallel worktree-agent prompts use RELATIVE paths only” — anchored in the convention memory.
    • New project convention #32 — “Strong gates absorb tooling failures” — locks the recovery posture.
  • Shared-secret bearer authentication for the coordinator (sprint 0.5-S3). Enables production deployment by gating every coord HTTP route on a bearer token. coordinator serve --shared-secret <hex> (or BORUNA_COORD_SECRET env var) and worker --shared-secret <hex> (same env var fallback) configure the symmetric secret. Mismatched or missing Authorization: Bearer header returns 401 + error_kind: coord.unauthorized. When unset, no auth is enforced — the pre-0.5-S3 loopback-only behavior is preserved for backwards compatibility, with a loud stderr warning when the coord binds to a non-loopback address without a secret.

    Generate a secret via openssl rand -hex 32. mTLS, per-worker keys, and OAuth integration deferred to 0.6.x — shared-secret covers the common operator case (single trusted cluster, per-deployment secret rotation).

    Auth applies to merged dashboard routes too — operators who want a public read-only dashboard with auth-gated mutations should run a separate boruna dashboard serve process. The coordinator’s merged listener is strictly all-or-nothing for auth.

    4 new CLI integration tests cover: missing bearer → 401, wrong bearer → 401, correct bearer → 200, no-secret legacy path → 200 (no regression).

  • Distributed retry policies (sprint 0.5-S5). Wires the existing RetryPolicy (max_attempts, on_transient, retry_on) through the wait driver so failed steps with retry budget transition Failed → Pending instead of permanent Failed. The coordinator stays dumb — all retry-decision logic lives in WorkflowRunner::advance_run_one_tick and a new persistence primitive RunCheckpointStore::requeue_failed_step_for_retry.

    The persistence primitive uses BEGIN IMMEDIATE + atomic status check inside the transaction; idempotent across concurrent wait clients (project convention §14). Returns a typed RequeueOutcome (Requeued { new_attempt_count }, NotFailed { current_status }, NotFound).

    AdvanceResult gains a newly_requeued: Vec<String> field (additive). The coordinator wait driver prints a distinct step <id>: requeued (retry) line for each requeued step before the generic transition print.

    Run-status derivation is updated: Failed is declared only when a step is Failed AND has no retry budget remaining. A Failed-with-budget step keeps the run Running and is requeued in the same tick.

    14 new orchestrator unit tests cover the retry pass: budget exhaustion, single-attempt rejection, error-class matching (retry_on vs. on_transient fallback), concurrent-wait race (idempotency), and policy-absent short-circuit. The pre-0.5-S5 wait limitation (“distributed retry not honored”) is now resolved; docs/design-coord-wait.md updated accordingly.

  • boruna fmt auto-formatter for .ax files (DX sprint, first item from the 0.2.x DX lane). Canonical pretty-printer that walks the existing compiler AST and emits formatted source.

    CLI: boruna fmt <file> rewrites in place; boruna fmt --check <file> exits 0 if the file is already formatted, exit 1 otherwise (CI gate). Exits 2 on parse errors so CI can distinguish “needs formatting” from “broken file”.

    Style decisions: 4-space indent, trailing comma on multi-line records and match arms, blank line between top-level decls, same-line opening braces.

    Known limitation (v1): comments are stripped — the lexer drops them before they reach the parser, so the current AST has no comment positions. A token-aware comment-preserving formatter is future work. v1 is still useful as a CI gate for generated/scaffolded code or code reviews where comments are preserved manually.

    3 golden-fixture tests, 1 idempotency roundtrip, 1 parse-failure error case, and 3 CLI integration tests (–check exit codes 0/1/2). New module tooling/src/format/ with format_source and check_source public APIs.

  • boruna coordinator wait <run-id> (sprint 0.5-S2f). Multi-wave workflow advancement for distributed runs. After workflow run --submit-only writes the first wave’s Pending checkpoints, coordinator wait polls runs.db, computes downstream-ready successors as workers complete steps, and writes Pending checkpoints for the next wave — repeating until the run reaches a terminal status.

    boruna coordinator serve --data-dir /var/lib/boruna &
    boruna worker run --coordinator http://127.0.0.1:8090 &
    boruna workflow run examples/workflows/document_processing \
        --submit-only --data-dir /var/lib/boruna
    # ↳ submitted run_id=...
    
    boruna coordinator wait <run-id> --data-dir /var/lib/boruna
    # ↳ polls every 500 ms, prints transitions per step,
    #    exits 0 on Completed / 1 on Failed
    

    Coordinator gains zero new logic — the “dumb transport” invariant from 0.5-S2c is preserved. All wave advancement is client-side. The wait driver is stateless: kill it at any point and re-invoke; the run continues from the persisted state.

    New flags on coordinator wait:

    • --poll-interval-ms <ms> (default 500, minimum 100; values below the floor are clamped with a warning).
    • --max-wait-secs <s> (default 0 = unlimited; useful for CI timeouts).

    Exit codes: 0 Completed, 1 Failed, 2 error (run not found, missing workflow_def, unsupported step kind in non- first wave), 3 --max-wait-secs exceeded.

    New persistence primitive RunCheckpointStore::insert_pending_step_if_absent(run_id, step_id) -> bool uses INSERT ... ON CONFLICT DO NOTHING so the wait client can safely write Pending checkpoints even when the coordinator is concurrently transitioning sibling steps. The legacy upsert_step_checkpoint (which hard-overwrites status on conflict) is unchanged; the new primitive is the race-safe variant for client-side advancement. Locked by insert_pending_step_if_absent_preserves_running_row.

    New field PersistedRunMetadata::workflow_def: Option<WorkflowDef> (with #[serde(default)] for back-compat). Embedded only when submit_only=true; capped at 1 MiB serialized JSON. In-process runs leave it None to keep metadata small.

    New WorkflowRunner::compute_ready_steps(def, status_map) (pure, deterministic-sort) and advance_run_one_tick(store, run_id) -> AdvanceResult (one polling tick).

    Tests: 14 new orchestrator unit tests covering the advance loop, the size cap, race-safe persistence, and idempotency. 4 new CLI integration tests: marquee multi-wave end-to-end, kill-and-resume, fail-on-bad-step, immediate-exit-on-already- completed.

    Known limitations (deferred to 0.5-S3+):

    • Retry policies in distributed mode — a step that fails is terminal; the wait driver exits with status 1 even if a retry policy would have succeeded in-process. Distributed retry is a future sprint.
    • Audit-chain coverage — the wait driver does NOT append WorkflowCompleted/WorkflowFailed events. Submit-only emits WorkflowStarted but the chain has no terminating entry for distributed runs. Auditors should check persisted runs.status directly.
    • Concurrent wait clients — multiple coordinator wait processes against the same run_id are safe (the race- safe primitive ensures idempotency) but not specifically tested as an integration scenario.
    • HTTP-based remote wait — the wait client requires filesystem access to --data-dir. A truly remote coordinator wait --coordinator <url> mode is a future sprint.
  • boruna workflow run --submit-only (sprint 0.5-S2e). The first end-to-end path for dispatching a real workflow through a coord+workers cluster. Submit-only mode: validates + computes the DAG, embeds source-step bodies in metadata_json.step_sources, inserts the run row + initial wave’s source-step Pending checkpoints, then exits before spawning thread workers. The cluster picks up the steps via existing claim/dispatch mechanisms.

    boruna coordinator serve --data-dir /var/lib/boruna &
    boruna worker run --coordinator http://127.0.0.1:8090 &
    boruna workflow run examples/workflows/llm_code_review \
        --submit-only --data-dir /var/lib/boruna
    

    Workflows using approval-gate / external-trigger steps in the first wave are rejected at submit time with a typed error (submit-only mode does not support ... in the first wave). Distributed mode for those features is deferred.

    Multi-wave automatic advancement is NOT done — operators monitor via the dashboard or boruna workflow show <run-id>. Wave loop integration becomes 0.5-S2f or later.

    Added field RunOptions::submit_only: bool and field PersistedRunMetadata::step_sources: BTreeMap<String, String> (with #[serde(default)] for back-compat). The WorkflowStarted audit event fires for submit-only runs matching the in-process run_persistent semantics.

    Tests: 3 new unit tests (insertion shape, metadata embedding, approval-gate rejection) + 1 new CLI integration test that runs boruna workflow run --submit-only against a real workflow.json + .ax file with a spawned coord+worker pair, asserts the step transitions through Pending → Running → Completed and the output_json matches the expected value.

  • Coordinator + dashboard listener-merge (sprint 0.5-S2d). The dashboard’s read-only routes (/, /runs/:id, /api/runs, /api/runs/:id) are now served on the same listener as the coordinator’s worker routes (/api/workers/..., /api/work/...). Operators get fleet visibility AND distributed dispatch from a single boruna coordinator serve invocation — one process, one port, one connection to runs.db.

    The merge is automatic — anyone running the coordinator gets the dashboard routes too. The standalone boruna dashboard serve keeps working unchanged for read-only deployments without the coordinator overhead.

    The coordinator’s bind_warning flows into the dashboard builder so the red HTML banner correctly fires when the coordinator is bound to a non-loopback address. Operators can’t accidentally expose the coordinator without the dashboard banner warning them.

    Refactor: dashboard::dashboard_routes(store, bind_warning) is now a pub route builder taking primitive args. The coordinator merges it onto its own router. Zero copy-paste; both the standalone dashboard and the coordinator use the same builder.

    3 new CLI integration tests cover the merged surface.

  • Coordinator background lease-expiry sweep (sprint 0.5-S2c). The coordinator now runs a tokio interval task that wakes up every --sweep-interval-ms (default 30 s) and calls expire_leases_and_requeue. Stale leases from worker crashes are now recovered without restarting the coordinator. Best-effort failure semantics: errors log + continue to the next tick.

    New CLI flag on boruna coordinator serve: --sweep-interval-ms <ms> (default 30000, minimum 100; values below the floor are clamped with a warning).

    New CLI integration tests:

    • coord_bg_sweep_requeues_expired_lease proves the sweep fires periodically and requeues stale leases without a coordinator restart.
    • worker_completes_two_step_linear_dag proves the protocol scales beyond a single step. (DAG advancement by the coordinator itself is deferred to 0.5-S2d; this test pre-populates both steps as Pending up front.)

    Architectural note documented in docs/design-coord-bg-sweep.md: in v0.5.x the coordinator is a “dumb transport” — it dispatches what’s in Pending and persists what completes. Wave advancement (deciding which step is Pending after a successful completion based on DAG dependencies) lives in the client. The boruna workflow run --coordinator <url> client mode ships in 0.5-S2d.

  • Coordinator/worker HTTP MVP (sprint 0.5-S2b). The HTTP layer over the persistence-layer state machine from 0.5-S2a. Two new CLI subcommands behind the serve feature flag:

    • boruna coordinator serve --data-dir <path> [--port 8090] [--bind 127.0.0.1] [--max-lease-ttl-ms 300000] [--poll-timeout-ms 30000]
    • boruna worker run --coordinator <url> [--worker-id <name>] [--lease-ttl-ms 300000] [--poll-timeout-ms 30000]

    Six HTTP routes per ADR 002: POST /api/workers/register, POST /api/workers/heartbeat, GET /api/work/claim (long-poll), POST /api/work/complete, POST /api/work/fail, POST /api/work/extend-lease. Every response carries protocol_version: 1. Worker-side: register → long-poll claim → compile + execute the step’s .ax source → POST result. Heartbeats every 10 s in a background task.

    Stable coord.* error_kind taxonomy, locked at this sprint’s ship: coord.lease_expired, coord.unknown_worker, coord.binary_mismatch, coord.invalid_request, coord.output_too_large, coord.step_not_found. The HTTP layer maps the persistence-layer outcome enums (from 0.5-S2a) 1:1 — no string-equality drift.

    Workers must match the coordinator’s capability_set_hash per ADR 002’s atomic-upgrade rule; mismatched workers get 409 + coord.binary_mismatch. Output payload size capped at 8 MiB per ADR 002; oversize bodies get 413 Payload Too Large from Axum’s DefaultBodyLimit.

    Workers parse policy via the strict validator from sprint 0.4-S15 (boruna_vm::policy_validate::parse), so workers reject the same shapes the CLI rejects with the same stable error_kind strings.

    On startup, the coordinator runs expire_leases_and_requeue to void any stale leases left over from a prior coordinator process (per ADR 002’s “coordinator restart = all leases void” rule).

    Loopback (127.0.0.1) by default. Non-loopback bind emits a loud stderr warning. No authentication — operators exposing the coordinator MUST front it with an auth-enforcing reverse proxy.

    Tests: 9 coordinator handler unit tests (route shapes, error_kind strings, status codes, lease-cap enforcement) + 4 worker unit tests (compile, execute, hash determinism, url-encoding) + 6 CLI integration tests including the flagship worker_kill_mid_step_lease_expires_then_reclaim regression that exercises the slow-but-not-dead worker race end-to-end at the wire level.

    New deps in the serve feature: reqwest 0.12 (json + rustls-tls, no openssl) for the worker’s HTTP client; uuid 1 for worker_id / session_token allocation.

    Not in this sprint (deferred to 0.5-S2c): workflow runner integration (boruna workflow run --coordinator <url>), wave-loop coordinator-side dispatcher, dashboard + coordinator listener-merge.

    See docs/design-coordinator-worker-http.md, docs/architecture-coordinator-worker-http.md, docs/test-plan-coordinator-worker-http.md.

  • Claim/lease persistence API (sprint 0.5-S2a). The persistence-layer half of ADR 002. Schema v3 adds three operational columns to step_checkpoints: worker_id (opaque worker handle), lease_expires_at (unix ms), and claim_id (monotonic per (run_id, step_id), CAS key for terminal-state transitions). Five new methods on RunCheckpointStore:

    • claim_step — atomic Pending → Running transition with incremented claim_id.
    • complete_step_cas — CAS-protected completion. Rejects late writes from expired-lease workers without changing persisted state.
    • fail_step_cas — CAS-protected terminal failure.
    • expire_leases_and_requeue — sweep expired leases back to Pending. Idempotent.
    • extend_lease_cas — push out the lease deadline, CAS-protected against the original claim_id.
  • New outcome enums: ClaimOutcome, TerminalOutcome, ExtendOutcome. Each carries a stable kind() -> &'static str per project convention #2 (claim.*, terminal.*, extend.*). These map to the wire-level coord.* error_kind strings the HTTP coordinator will lock in 0.5-S2b.

  • Schema v2 → v3 migration via the existing migration runner. Idempotent — re-opens are no-ops; fresh databases get the full v3 schema directly from SCHEMA_V1_SQL.

  • 32 new persistence tests including the load-bearing slow_worker_race_late_completion_rejected regression that exercises the slow-but-not-dead worker race the ADR’s adversarial review caught: claim → expire → reclaim → original worker’s late completion → LeaseExpired rejection → row state unchanged. If this test ever fails, the state machine is broken.

  • The single-process WorkflowRunner path is unchanged. upsert_step_checkpoint does not write the new columns; they stay at their defaults (None, 0) for steps that flow through the in-process scheduler.

Decided

  • ADR 002 — Distributed step execution. The 0.5.0 (“Scale”) cycle’s foundational architectural decision. Distributed mode uses an embedded HTTP coordinator + lightweight HTTP workers, all behind the existing serve feature flag. The coordinator remains the only writer of runs.db; workers long-poll for claimable steps and report results via JSON over HTTP. Lease- based claim with re-dispatch on expiry handles worker crashes. Determinism is preserved: which worker ran a step is operational state and never enters the audit/replay pipeline. The single-process path (boruna workflow run/resume) keeps working unchanged. Considered alternatives — shared-filesystem SQLite, external queue (Redis/RMQ/SQS), gRPC — were rejected for footgun risk, deployment-simplicity violation, and marginal benefit respectively. Implementation in 0.5-S2. See docs/adr/002-distributed-step-execution.md.

0.4.0 — 2026-04-27

The operations release. Twelve sprints (0.4-S5 through 0.4-S16) ship the production-readiness layer on top of 0.3.0’s durability work: distributed-tracing observability, streaming progress, multi-pause-per-level wave loops, per-error-class retry classification, hash-chained audit decisions and lifecycle events, post-hoc evidence-bundle creation, Prometheus metrics, multi-provider LLM dispatch, multi-environment data separation, strict-validated policy-as-code, and a read-only HTTP dashboard.

Added

  • Workflow dashboard (sprint 0.4-S16). New boruna dashboard serve subcommand exposes a read-only HTTP view of runs.db so operators can triage at a glance without dropping into sqlite3. Loopback (127.0.0.1) by default; --bind 0.0.0.0 is allowed but shouts a loud warning on stderr AND renders a red banner in the HTML, because the dashboard ships with no authentication.

    cargo build --release -p boruna-cli --features serve
    boruna dashboard serve --data-dir /var/lib/boruna
    

    Routes: GET / (HTML index), GET /runs/:id (HTML detail), GET /api/runs (JSON list), GET /api/runs/:id (JSON detail). Zero mutation routes — POST/PUT/DELETE/PATCH to any path return 405. Multi-env aware: when --env is set, the dashboard reads <data-dir>/<env>/runs.db per the 0.4-S14 contract.

    Builds behind the existing serve feature flag (already used by boruna serve for framework apps). Reuses the workspace axum 0.8 + tokio deps.

  • boruna_orchestrator::persistence::{RunRow, RunRecord, RunOperational, StepCheckpoint} now derive Serialize so read-only consumers can render rows directly. (Not Deserialize — there’s no scenario where a dashboard consumer should be reconstructing a row.)

  • 18 new unit tests in dashboard::tests covering every handler, HTML escaping (XSS regression), bind-warning banner, 404, and the date-format helper. 8 new CLI integration tests in crates/llmvm-cli/tests/cli_dashboard.rs covering end-to-end HTTP behavior, the read-only contract (POST → 405), and CLI error paths (missing data-dir, invalid bind address).

  • New CI steps to build and test the serve feature (cargo build/test/clippy -p boruna-cli --features serve).

  • New reference doc docs/reference/dashboard.md covering build, run, security posture, routes, stability tier.

  • Policy management as code (sprint 0.4-S15). Operators now treat --policy files as versioned, validated, code-reviewable artifacts. Two new CLI subcommands:

    • boruna policy validate <file> [--json] — strict-validate a policy file. Designed as a CI gate. Exits 0 on ok, 2 on validation error, 1 on file IO error.
    • boruna policy show <file> — validate then print the effective policy (default behavior, denormalized rule list, net_policy bounds). Plus a new MCP tool boruna_policy_validate(policy_json) that runs the same validator. The CLI, MCP, and boruna run --policy <file> paths now share one parser — passing validate but failing run is structurally impossible.
  • New boruna_vm::policy_validate::{parse, parse_file, PolicyParseError, POLICY_SCHEMA_VERSION}. The validator enforces:

    • schema_version ∈ {1} (other values rejected — locks the contract for forwards-compat).
    • Top-level / net_policy / per-rule fields are an allow-list — unknown fields rejected (policy.unknown_field). Closes the silent-default footgun where "default_alow": true parsed as default_allow: false.
    • rules keys must be canonical capability names. Aliases ("net", "db", …) rejected with a hint to the canonical name ("net.fetch", "db.query"). Aliases used to silently no-op at gateway-check time.
    • net_policy.max_response_bytes > 0, timeout_ms > 0, allowed_methods ⊆ {GET, POST, PUT, DELETE, PATCH, HEAD, OPTIONS} (canonical upper-case; lower-case rejected).
  • Stable error_kind taxonomy — locked per project convention #2: policy.io_error, policy.parse_error, policy.unknown_schema_version, policy.unknown_field, policy.invalid_capability, policy.invalid_net_policy. Future validators can add new kinds; existing kinds never rename.

  • 26 unit tests in boruna_vm::policy_validate + 11 CLI integration tests in crates/llmvm-cli/tests/cli_policy.rs + 7 MCP tests + 3 protocol_version regression tests.

  • Updated docs/reference/policy-schema.md with the strict-validator rules, error_kind taxonomy, and CLI tooling examples. Design rationale in docs/design-policy-as-code.md.

Fixed

  • Policy::default() now produces schema_version: 1 (matching what the lenient deserializer’s #[serde(default = "...")] produces for an empty input). The derived default leaked schema_version: 0 into round-trips — invisible until the 0.4-S15 strict validator surfaced it. Affects Policy::deny_all() and any caller that started from Policy::default().

  • Multi-environment support (sprint 0.4-S14). New global --env <name> flag (also from BORUNA_ENV env var). When set:

    • --data-dir is namespaced to <data-dir>/<env>/ so each environment has its own runs.db, audit chains, and evidence bundles.
    • Every Prometheus metric gains an env="<env>" label so dashboards can filter / group by environment.
    boruna --env staging workflow run wf --data-dir /var/lib/boruna ...
    boruna --env prod workflow run wf --data-dir /var/lib/boruna ...
    # → /var/lib/boruna/staging/ and /var/lib/boruna/prod/ stay separate
    

    Operators get dev/staging/prod separation without external orchestration. Per-env policy is supplied via --policy per call.

  • New boruna_orchestrator::metrics::format_prometheus_with_env variant. Backward compatible: format_prometheus(snap) continues to produce env-less output (calls format_prometheus_with_env(snap, None) internally).

  • New CLI helper validate_env_name rejects names with characters outside [a-zA-Z0-9_-] (length 1-64). Protects against path traversal (--env ../../etc/passwd is rejected at the boundary) and broken Prometheus labels.

  • 4 new tests in metrics: env label added to every series, env-less output is byte-identical to legacy, env label escapes, end-to-end BORUNA_ENV round-trip via the export entry.

Backward compatibility

When --env and BORUNA_ENV are both unset, behavior is exactly as before: data goes to <data-dir>/, metrics carry no env label. Operators upgrading from 0.4-S13 see no change unless they opt in.

  • LlmRouterHandler — multi-provider LLM dispatch helper (sprint 0.4-S13). Direct extension of the BYOH decision in 0.3-S8. Integrators with multiple LLM providers (OpenAI + Anthropic + local Ollama / vLLM) no longer need to write their own dispatch logic — the router takes a registry of provider handlers and routes each Capability::LlmCall based on a provider/model prefix in args[1]:

    #![allow(unused)]
    fn main() {
    let mut providers: BTreeMap<String, Box<dyn CapabilityHandler>> = BTreeMap::new();
    providers.insert("openai".into(), Box::new(my_openai_handler));
    providers.insert("anthropic".into(), Box::new(my_anthropic_handler));
    let router = LlmRouterHandler::new(providers, Box::new(MockHandler));
    }

    .ax callers then write llm_call("Summarize:", "openai/gpt-4") — the prefix selects the provider; the full model string (including the prefix) is forwarded unchanged so providers can use it for internal tagging.

  • The router is pure routing logic — Boruna still ships zero provider HTTP code. Each provider’s handler implementation, authentication, and response parsing belong to the integrator per the BYOH model.

  • Non-LLM capability calls pass through to a fallback handler so the router composes with the existing StepInputHandler / MockHandler / HttpHandler stack.

  • Typed errors for: missing model arg, non-string model arg, malformed model string (no /), empty provider prefix, unknown provider (error message includes the registered providers list).

  • Late-registration support via add_provider(name, handler) returning the previously-registered handler.

  • 11 unit tests covering routing, args forwarding, error variants, fallback delegation, late registration, and deterministic registered_providers ordering.

  • Updated docs/guides/llm-integration.md with a new section walking through the router setup.

  • Prometheus metrics export CLI (sprint 0.4-S12). New boruna metrics export --data-dir <DIR> command reads the persistent run store and writes Prometheus text format to stdout. Operators integrate via cron + node_exporter’s textfile collector — the canonical Prometheus pattern for batch tools:

    */30 * * * * boruna metrics export --data-dir /var/lib/boruna \
                    > /var/lib/node_exporter/textfile_collector/boruna.prom
    

    Architectural decision documented in docs/design-prometheus-metrics.md: CLI-pulled (not embedded HTTP) to align with Boruna’s CLI-only philosophy locked in 0.3-S15 (BYOH webhook pattern). No new long-running daemon process.

  • Three metric families:

    • boruna_workflow_runs_total{workflow,status} — counter of runs by terminal/transient status.
    • boruna_workflow_runs_in_flight{workflow} — gauge of running or paused runs.
    • boruna_workflow_step_completions_total{workflow,step,status} — counter of step terminal transitions (completed / failed).
  • New boruna_orchestrator::metrics module with compute_snapshot, format_prometheus, and export public entries. The snapshot is pure data so future exporters (JSON dashboard endpoint, etc.) can reuse it without re-querying the store.

  • 8 unit tests covering: empty store emits HELP+TYPE only, aggregation by workflow/status, in-flight counting, terminal step transitions only (no Pending/Running noise), output is valid Prometheus textfile format with HELP/TYPE preceding data, determinism (BTreeMap iteration locked), label escaping (backslashes, quotes, newlines per the exposition spec), end-to-end realistic run set.

Counter semantics caveat

Counters are computed from current store state at sample time, not maintained as deltas. If old runs are pruned from the DB, the _total will decrease — Prometheus normally treats this as a counter reset and handles it gracefully via rate(). Operators running frequent pruning should be aware of this contract.

  • Full lifecycle audit events (sprint 0.4-S11). Closes the audit theme for 0.4.0. The audit chain now captures the complete run lifecycle, not just operator decisions:
    • WorkflowStarted { workflow_hash, policy_hash } — appended at execute_after_insert’s top, immediately after the run row inserts.
    • StepCompleted { step_id, output_hash, duration_ms } — appended after each step’s terminal Completed checkpoint write.
    • StepFailed { step_id, error } — appended after each step’s terminal Failed checkpoint write (including panic-failed workers in the concurrent path).
    • WorkflowCompleted { result_hash, total_duration_ms } — appended at terminal status only (Completed/Failed). Resume’s terminal exit also appends it. Pause states leave the chain open for the next resume to extend.
  • New append_audit_event(store, run_id, event) helper using the same CAS-retry pattern as record_approval_decision / record_external_trigger. Lifecycle appends are best-effort: a CAS budget exhaustion logs a warning and continues. Missed audit events are operationally annoying (chain has fewer step events than checkpoints) but never fail the run — the chain entries that DID commit remain valid, and an auditor at verify time sees the gap explicitly.
  • StepStarted events are deliberately NOT emitted — the checkpoint’s started_at_ms already captures per-step start operationally, and emitting an event-per-start would double the CAS-write count for limited compliance value.
  • 2 new tests in tests::evidence_bundle: lifecycle_events_emitted_in_order_for_multi_step_run (4-entry chain in topological order: Started → 2× StepCompleted → Completed) and step_failed_event_emitted_on_runtime_error (chain captures the failed step + error message).
  • 7 existing audit_decisions / evidence_bundle tests updated to match the new chain shape (lifecycle events + decisions).

Audit theme summary

Across 0.4-S9 (decisions), 0.4-S10 (bundle creation), and 0.4-S11 (lifecycle events), the audit story is now end-to-end complete: every persistent run produces a hash-chained audit log of all lifecycle transitions and operator actions, the chain is persisted atomically with the corresponding state changes, and boruna evidence create <run-id> packages it with all reproducibility artifacts for downstream verification via boruna evidence verify.

Performance impact

For a workflow with N steps, the chain now requires roughly N+2 additional CAS-protected metadata writes (1 WorkflowStarted, N StepCompleted/Failed, 1 WorkflowCompleted). Each write is a single SQLite UPDATE with a small JSON blob. For typical workflows this is operationally negligible. High-throughput integrators can disable lifecycle audit by deferring this sprint’s wiring (no disable flag ships in this sprint — file an issue if needed).

  • boruna evidence create <run-id> (sprint 0.4-S10). Builds an evidence bundle from a persisted run by reading the run’s metadata, step checkpoints, and hash-chained audit log. Closes the audit-evidence loop end-to-end:
    $ boruna workflow run wf --data-dir .data --policy allow-all
    $ boruna workflow approve <run-id> <step-id> --data-dir .data
    $ boruna workflow resume <run-id> --data-dir .data
    $ boruna evidence create <run-id> --output-dir bundles --data-dir .data
    $ boruna evidence verify bundles/<run-id>      # VALID
    
  • New boruna_orchestrator::workflow::create_bundle(data_dir, run_id, output_dir) public entry. Reads workflow.json from the run’s recorded workflow_dir, policy from the persisted policy_json column, per-step outputs from step_checkpoints.output_json, and the full audit chain from metadata.audit_log (sprint 0.4-S9). Builds an EvidenceBundleBuilder, finalizes, returns the BundleManifest.
  • 6 new tests in tests::evidence_bundle: complete artifact for a completed run, audit chain round-trip via JSON, end-to-end verify_bundle() passes on the produced bundle, trigger payload hash matches the synthesized step output_hash, unknown run id returns typed RunNotFound, runs without decisions produce an empty chain whose audit_log_hash is the all-zeros sentinel.

Post-hoc bundle creation

The runner does NOT auto-create bundles during execution — the hot path stays free of bundle I/O. Operators trigger bundle creation explicitly when needed (e.g., a compliance request months after the run completed). Same model as the rest of the audit subsystem: operator-driven, not runner-driven.

  • Audit-log integration of approval / trigger decisions (sprint 0.4-S9). Closes a 0.3.0 carried-forward debt. Operator actions (approval grants/denials, external trigger events) now produce hash-chained audit entries, persisted as metadata.audit_log and written atomically with the operator-facing decision via the existing CAS-protected metadata writes.
  • New AuditEvent::ExternalTriggerReceived { step_id, payload_hash } variant. The payload_hash matches the synthesized step output_hash (since the trigger payload becomes the step’s output value), so the chain links to the replay-verified output. Payload itself is hashed rather than logged verbatim — webhook bodies may contain operator PII.
  • New AuditLog::from_entries(Vec<AuditEntry>) -> Self and AuditLog::into_entries(self) -> Vec<AuditEntry> for round- tripping the chain through a containing struct (e.g. the run’s persisted metadata) without re-serializing to JSON.
  • 7 new tests covering: approval-grant / approval-reject append the right event, trigger appends with payload_hash equal to output_hash, multi-decision chain integrity (prev_hash chains), legacy 0.3.x metadata round-trip without audit_log field, audit log persists unchanged across resume, first decision after legacy metadata starts a fresh genesis chain.
  • Design doc: docs/design-audit-decision-events.md.

Tamper-evidence vs replay-verification

The audit chain’s prev_hash linkage is tamper-evident — any post-hoc mutation (direct sqlite3 surgery, bit-flip in storage) surfaces when an auditor calls AuditLog::verify(). The chain is not processed by the run’s deterministic-execution replay pipeline; replay verifies per-step output_hash, not the operator-action chain. Documented prominently in the PersistedRunMetadata.audit_log doc-comment to prevent confusion with the replay-verification subsystem.

Backward compatibility

A 0.3.x metadata blob with no audit_log field deserializes via #[serde(default)] to Vec::new(). The first decision recorded by a 0.4-S9 binary on a 0.3.x run starts a fresh genesis chain (sequence=0, prev_hash=“0”*64). Locked by first_decision_after_legacy_metadata_starts_chain_at_sequence_zero.

What this sprint does NOT ship

  • Full lifecycle audit events (WorkflowStarted, StepStarted, StepCompleted, etc.) — separately scheduled. This sprint surgically closes the operator-action audit gap without touching the per-step hot path.

  • Audit log in evidence bundles — EvidenceBundleBuilder::finalize already accepts an AuditLog parameter; wiring the in-metadata log into bundle construction is a small follow-on sprint.

  • Operator identity capture — no auth subsystem yet. The approver field is empty string until a future identity sprint wires real auth. The field IS captured in the hash chain regardless so a future upgrade can fill it in without re-keying past entries.

  • Per-error-class retry classification (sprint 0.4-S8). The RetryPolicy schema gains an explicit retry_on: Vec<String> allowlist alongside the legacy binary on_transient gate. Operators who want “retry on transient timeouts but NOT on auth errors or bad code” now express it directly:

    "retry": {
      "max_attempts": 3,
      "on_transient": false,
      "retry_on": ["wall_time_exceeded", "io_error"]
    }
    
  • New error_class taxonomy with stable string constants: wall_time_exceeded, step_limit_exceeded, capability_denied, capability_budget_exceeded, compile_error, runtime_error, io_error, input_resolution. Forward-compatible — new classes add without breaking existing policies.

  • New classify_vm_error(&VmError) -> &'static str maps every VM error variant to its taxonomy class. Catch-all is runtime_error (assertions, type errors, OOB, divisions, stack errors, bytecode errors all surface here).

  • should_retry_class(policy, class) -> bool — central decision function. Resolution order: no policy / max_attempts ≤ 1 → false; non-empty retry_on → match in list; empty → fall back to on_transient.

  • retry_with_backoff short-circuits on non-retry-eligible failures rather than running through the full backoff schedule. A compile error no longer waits 100+200+400ms before giving up.

  • 17 new tests covering classification mappings, allowlist semantics, legacy fallback, unknown-class-ignored, retry_on takes precedence over on_transient=false, and serde round-trip for legacy 0.3.x workflow.json files (no retry_on field).

Backward compatibility

  • A 0.3.x workflow.json with retry: {max_attempts, on_transient} (no retry_on field) deserializes with retry_on = vec![] via #[serde(default)]. The empty allowlist falls back to the legacy on_transient gate, so prior behavior is exactly preserved.

  • class strings are case-sensitive. Use the lowercase snake_case forms documented in error_class::*. Unknown strings (typos like "transient_netwrok") are silently ignored — they never match a real failure class, so the policy behaves as if the typo were absent (conservative-by-default).

  • Wave-loop multi-pause-per-level (sprint 0.4-S7). The concurrent execution path (--concurrency >= 2) now pauses ALL pause-steps in the same DAG level in a single execution pass — previously only the first was processed and remaining pauses were silently deferred to subsequent resumes. Enables “wait for payment AND fraud-check” webhook fan-in patterns where multiple external_trigger (or approval_gate) steps depend on a shared upstream and a downstream step depends on all of them. Each pause persists its own checkpoint and (for trigger steps) mints its own distinct token. The resume sentinel pass advances each pause independently as its decision/event arrives.

  • New persist_one_pause helper isolates per-pause persistence errors. If one pause’s acquire_trigger_token or upsert_step_checkpoint fails (transient /dev/urandom error, CAS retry exhaustion, disk error), the loop logs a warning and continues to the next pause. The run is marked Paused on the pauses that DID commit, leaving operators with a recoverable state. The next resume’s wave loop is idempotent — acquire_trigger_token reuses existing tokens and upsert_step_checkpoint is re-write-safe — so the failed pauses retry cleanly. Reviewed in 0.4-S7 — earlier draft propagated the first per-pause error, terminally-failing the run and stranding pause #1’s token with no recovery path.

  • 5 new tests in tests::multi_pause_per_wave: 2-trigger parallel pause, partial trigger fire keeps other paused, full trigger fire advances downstream, mixed approval+trigger pauses, partial-pause failure recovery via direct-SQL state injection.

Asymmetry note

The sequential execution path (--concurrency 1) is unchanged: it processes one step at a time and serializes parallel pauses across multiple resumes. Operators expecting AND-fan-in webhook patterns must use --concurrency 2 or higher.

  • Streaming progress notifications from boruna_run (sprint 0.4-S6, closes #4). When the MCP caller supplies a progressToken in the request _meta field (per the MCP spec), the server emits notifications/progress events with the cumulative VM step count every 100k opcodes. Long-running scripts no longer block the calling agent’s UI behind a single final result blob. Backward compatible: callers without a progressToken see the legacy synchronous behavior unchanged.
  • New Vm::start_timer() method — initializes the wall-clock timer used by max_wall_ms budgets. Callers driving the VM through execute_bounded should call it before set_entry_function to match Vm::run’s timing contract (the entry-frame allocation counts toward the budget).
  • New Vm::set_in_actor_context(bool) flag — replaces the prior budget.is_some() heuristic for distinguishing actor-system scheduling from standalone bounded execution. Op::ReceiveMsg on an empty mailbox now blocks (rewind IP + MailboxEmpty) only when the flag is set; standalone bounded loops fall through with Value::Unit, matching Vm::run’s legacy semantics. Reviewed in 0.4-S6 — without this fix, the streaming-progress and non-streaming paths of boruna_run would diverge for any program emitting Op::ReceiveMsg outside an actor system.
  • ActorSystem::run sets in_actor_context = true on the root and every spawned child VM.

0.3.0 — 2026-04-26

Theme: Real-use durability. 0.3.0 makes Boruna usable for long-running, durable, production workflows. Persistent state survives process restarts; concurrent steps fan out within waves; transient failures retry with backoff; webhook-driven steps wait for external events. The full sprint stack (0.3-S2a through 0.3-S16) closes every big-rock theme on the original 0.3.0 plan and adds review- driven safety work.

Added

  • Persistent workflow state (sprints 0.3-S2a/S2b/S3/S6). Crash-resumable runs via SQLite-backed checkpoint store with BEGIN IMMEDIATE atomicity, f_FULLFSYNC on macOS for durability, and a --data-dir flag on boruna workflow run / resume.
  • Approval-gate operator UX (sprint 0.3-S2c). New kind: "approval_gate" step type pauses the run; operators advance via boruna workflow approve <run-id> <step-id> / boruna workflow reject with optional reason. Decisions persisted in run metadata.
  • Concurrent step execution within waves (sprint 0.3-S4). --concurrency N on run / resume parallelizes steps at the same DAG topological level. Determinism preserved: same output_hash regardless of concurrency.
  • Step retry policies (sprint 0.3-S5). Configurable per-step retry with exponential backoff (100ms × 2^N capped at 5s) for transient failures.
  • Idempotent invocation (sprints 0.3-S7 + 0.3-S10). --skip-if-running flag for cron-driven scheduling. Atomic skip-if-in-flight check + insert in a single transaction closes the prior race window.
  • LLM handler decision: Bring Your Own Handler (sprint 0.3-S8). No default LLM handler ships in core; integrators wire their provider via the CapabilityHandler trait. Reference OpenAI handler + integration contract in docs/guides/llm-integration.md.
  • Workflow versioning for CI/CD safety (sprint 0.3-S9). --expect-workflow-hash flag refuses runs / resumes when the on-disk definition’s hash doesn’t match.
  • Per-step attempt_count column (sprints 0.3-S11/S12/S13) with the project’s first schema migration (v1→v2) via column_exists + if v < N pattern. boruna workflow show surfaces the column. Sequential failure path persists actual count.
  • Workflow step output piping via step_input builtin (sprint 0.3-S14). let received: String = step_input("name") returns the JSON-encoded upstream output. New Capability::StepInput (id=10). Both sequential and concurrent paths resolve inputs coordinator- side. Unknown input names error with the declared list (review- driven).
  • Async step execution via external trigger CLI (sprint 0.3-S15). New external_trigger step kind for webhook-driven workflows. boruna workflow trigger <run-id> <step-id> --token <X> --payload <json> records the payload as the step’s output value. 32-hex-char tokens from /dev/urandom (no fallback) prevent accidental cross-step triggers. Constant-time validation; webhook- replay rejected by StepAlreadyTriggered. Boruna stays a CLI tool — no in-binary HTTP server.
  • Real HTTP handler with SSRF protection (added Feb 2026). Feature-gated http builds enable real network calls via --live. NetPolicy allowed_domains / methods / byte limits / timeout. Rejects private IPs, localhost, non-http schemes.
  • 23 new typed errors covering approval-gate, trigger-gate, run-not- resumable, step-not-found, hash-mismatch, and CAS-budget-exhausted states.

Fixed

  • Trigger-flow TOCTOU race (sprint 0.3-S16). The 0.3-S15 trigger flow split metadata writes (CAS) and step-checkpoint transitions (resume sentinel pass) across two separate SQL transactions. A concurrent boruna workflow resume calling mark_step_running_clearing_output between the trigger function’s metadata-CAS and the next resume’s sentinel pass could leave the payload silently logged-and-discarded. Fixed by wrapping the metadata CAS and the checkpoint transition in a single BEGIN IMMEDIATE SQL transaction (new RunCheckpointStore::commit_external_trigger). SQLite’s write-locked transaction blocks concurrent writers, making the checkpoint state read inside the transaction authoritative.
  • New TriggerCommitOutcome enum (Committed | MetadataChanged | CheckpointStateMismatch { current_status }) for callers that need to distinguish CAS-retry-eligible races from operator-error states.
  • Resume sentinel pass remains in place as a defensive recovery for legacy 0.3-S15-format DBs (metadata.triggers populated with non-empty payload but checkpoint still in awaiting_external_event). New forward-compat test confirms the upgrade path.
  • 5 new persistence-layer unit tests + 3 new runner-level integration tests cover the atomic-commit outcomes and the legacy upgrade scenario.

Added

  • Async step execution via external trigger CLI (sprint 0.3-S15). New external_trigger step kind for webhook-driven workflows. The runner pauses at the gate; an operator (or webhook receiver) advances it with boruna workflow trigger <run-id> <step-id> --token <X> --payload <json>, and the payload becomes the step’s output value (visible to downstream steps via step_input).
    "webhook": {
      "kind": "external_trigger",
      "description": "Stripe payment.succeeded webhook",
      "depends_on": ["init"]
    }
    
    Pause-time prints a 32-hex-char trigger token (16 bytes from /dev/urandom); the CLI rejects mismatching tokens to prevent accidental cross-step triggers from a misrouted webhook. Boruna stays a CLI tool — no in-binary HTTP server. The operator’s webhook receiver bridges to the CLI.
  • New StepKind::ExternalTrigger { description } variant on workflow step definitions; new StepStatus::AwaitingExternalEvent (persisted as "awaiting_external_event").
  • Public entry boruna_orchestrator::workflow::record_external_trigger for programmatic embedders. Validates the run/step/state, validates the operator-supplied token in constant time, refuses replays of already-triggered steps (StepAlreadyTriggered { prior_triggered_at_ms }), and writes the payload via compare-and-swap.
  • Resume sentinel pass advances paused trigger steps when a payload is recorded (mirrors the approval-decision pattern from sprint 0.3-S2c). The payload is stored as Value::String(payload); the audit hash chain captures the synthesized output_hash.
  • Five new typed errors: NotAnExternalTriggerStep, StepNotAtExternalTriggerGate, InvalidTriggerToken, StepAlreadyTriggered, plus an empty-payload Validation guard.
  • Ephemeral runs reject external_trigger steps upfront (review-driven). WorkflowRunner::run (no persistence) refuses workflows that contain trigger steps with a typed Validation error — earlier draft caught this at step-entry time, which silently allowed prior steps to execute before the typed error surfaced.
  • Trigger token reuse across resume (review-driven). The token is acquired via acquire_trigger_token: if a previously-persisted token exists for the step, it’s returned verbatim. Earlier draft generated a fresh token on every pause entry while persist-trigger-token’s “leave existing” branch kept the original; the printed value would silently disconnect from the validated value, and operators copying the just-printed token would get InvalidTriggerToken.
  • No fallback for entropy failure (review-driven). If /dev/urandom cannot be read, generate_trigger_token returns Err. Earlier draft fell back to a SystemTime + pid + counter hash, which gave low-entropy observer-predictable tokens silently.
  • Workflow step output piping via step_input (sprint 0.3-S14). New built-in function in .ax:
    let received: String = step_input("msg")
    
    Returns the JSON-encoded upstream output for the named input (declared in workflow.json’s inputs: { msg: "upstream.result" }). Steps that need typed access parse the JSON inline. Determinism preserved: same inputs → same per-step output_hash regardless of concurrency level.
  • New Capability::StepInput (id=10, name=“step.input”, version=“1”). Bumps capability_set_hash — additive surface change. Integrators using the prior hash for cache keys MUST invalidate. Old: sha256:b0ca1793.... New: sha256:980d017d....
  • Compiler treats step_input(name) as a builtin (typeck arity 1; codegen emits Op::CapCall(StepInput, 1)). Auto-adds Capability::StepInput to the calling function’s capability set so the VM’s runtime function-cap check passes.
  • New boruna_vm::capability_gateway::StepInputHandler — wraps an inner handler and intercepts step.input calls. Composes with both MockHandler and BYOH live handlers (sprint 0.3-S8).
  • WorkflowRunner::build_step_policy auto-allows step.input when the operator’s policy is silent on it. entry().or_insert() preserves explicit denies for hardened workflows.
  • Both sequential and concurrent execution paths resolve inputs coordinator-side and pass the snapshot to workers — workers hold no DataStore reference.
  • Unknown input names error (review-driven, project-conventions §1). step_input("undeclared_name") returns a typed runtime error with the declared list for triage, instead of silently returning empty data and corrupting downstream output.

Fixed

  • Sequential failure path persists actual attempt_count (sprint 0.3-S13, closes carried-forward limitation from 0.3-S11). Prior to this, the sequential execute_steps failure branch defaulted to attempt_count=1 even after retry exhaustion — so a step configured with max_attempts: 3 that exhausted all 3 attempts showed up as attempt_count=1 in the persisted SQL row and on workflow show. The error message correctly said “failed after 3 attempts” but the column lied. Fix: execute_source_step now returns Result<StepResult, (WorkflowRunError, u32)> carrying the count on both branches; the caller threads it through to the Failed checkpoint upsert. Concurrent path was already correct.

Added

  • workflow show surfaces attempt_count (sprint 0.3-S12). Plain mode adds an ATTEMPTS column to the steps table; --json mode adds attempt_count to each step’s object. Closes the operator-visibility loop opened by 0.3-S11 — operators triaging flaky steps no longer need to query SQLite directly.

  • step_checkpoints.attempt_count column (sprint 0.3-S11). Tracks the number of attempts each step took to reach its terminal state — 1 for first-try success or single-attempt failure; >1 when the retry policy fired (sprint 0.3-S5). Operational only — wall-clock-keyed (depends on whether transient failures happened); never feeds an audit hash. Surfaced on StepResult, StepCheckpoint, and persisted in the SQL store. First real schema migration: bumps SCHEMA_VERSION to 2; existing v1 databases are migrated additively via ALTER TABLE ADD COLUMN with DEFAULT 1 (no rewrite, instant). The migration runner is idempotent — fresh databases (where the canonical creation script already includes the column) skip the ALTER.

  • New library API:

    • RetryPolicy-aware retry_with_backoff now returns Result<(T, u32), (E, u32)> so callers can persist the actual attempt count alongside success or failure.
    • compile_and_run_step_with_retry returns (Value, u32) / (WorkflowRunError, u32) — same change in the runner-level wrapper.
    • StepResult.attempt_count: u32 (defaults to 1 for back-compat on older serialized JSON).
    • StepCheckpoint.attempt_count: u32 matches the SQL column.
    • persistence::SCHEMA_V1_TO_V2_SQL and persistence::column_exists helpers exposed within the crate.

Fixed

  • --skip-if-running race window closed (sprint 0.3-S10, carried-forward debt from 0.3-S7). Prior implementation’s two-call flow (find_in_flight_runs then run_persistent) let two concurrent processes both pass the in-flight check and both insert new run rows. Now folded into a single BEGIN IMMEDIATE SQL transaction via the new RunCheckpointStore::insert_run_with_derived_id_skip_if_in_flight method: at most one of N concurrent invocations inserts; the rest cleanly Skip. Locked by an 8-thread regression test that asserts exactly 1 Inserted + 7 Skipped outcomes. New library API: WorkflowRunner::run_persistent_or_skip returning Option<WorkflowRunResult> (Some = ran, None = skipped). The CLI flow now uses this atomic path under --skip-if-running.

Added

  • --expect-workflow-hash <HEX> on boruna workflow run and boruna workflow resume (sprint 0.3-S9). CI/CD safety primitive that refuses to start (or resume) if the on-disk workflow def’s workflow_hash doesn’t match the operator-supplied expected value. Catches accidental edits, malicious mutation, and stale- checkout-vs-config drift before any side effect.
  • --print-hash on boruna workflow validate. After validation succeeds, emits workflow_hash=<64-char hex> on its own stdout line — cut-friendly for shell pipelines:
    HASH=$(boruna workflow validate ./wf --print-hash | grep ^workflow_hash | cut -d= -f2)
    boruna workflow run ./wf --expect-workflow-hash $HASH ...
    
    Hash comparison is case-insensitive + whitespace-trim-tolerant so operators can paste from any source.
  • Note: the hash covers the workflow.json structure only — .ax step source changes do NOT affect the hash. For full-source coverage operators should hash the workflow_dir tree at the filesystem layer.

Decided

  • LLM live handler model: Bring Your Own Handler (BYOH) (sprint 0.3-S8). Boruna does NOT ship a default LLM handler in core. Integrators implement the CapabilityHandler trait against their provider of choice (OpenAI, Anthropic, vLLM, Ollama, custom routers) and pass it to CapabilityGateway::with_handler at workflow run time. Rationale: provider churn shouldn’t destabilize Boruna releases; API-key management belongs in the integrator’s application; production integrators (FleetQ et al.) already have their own LLM clients. New guide: docs/guides/llm-integration.md covers the contract, provider variants, determinism notes, and testing patterns. Reference handler at examples/llm_handlers/openai/. Closes the open question carried since the original 0.3.0 plan; docs/roadmap.md and docs/limitations.md updated accordingly.

Added

  • boruna workflow run --skip-if-running (sprint 0.3-S7). Idempotent invocation primitive for cron-driven scheduled workflows. Before launching a new run, queries the persistent store for any in-flight (Running or Paused) run of the same workflow. If found, exits 0 cleanly with a stderr message identifying the prior run. Designed for the cron pattern:
    0 2 * * * boruna workflow run /path/to/wf \
              --skip-if-running --data-dir /var/lib/boruna
    
    Without this flag, overlapping invocations could race on the same outputs/ directory and double-bill external API calls. Persistent path only; rejected at parse with --ephemeral.
  • New library API: boruna_orchestrator::workflow::find_in_flight_runs(data_dir, def), boruna_orchestrator::persistence::RunCheckpointStore::list_in_flight_runs_for_workflow.

Fixed

  • Power-loss durability for DataStore::store_output (sprint 0.3-S6, closes H1/C3 deferral from 0.3-S3). After tempfile::persist, the parent directory is now opened and fsynced so the rename’s directory entry is journaled to stable storage. Without this, POSIX permits the dirent to be lost on power loss even though the file’s data blocks were flushed. On macOS uses fcntl(F_FULLFSYNC) for both file and directory syncs (review-driven 0.3-S6 finding) — plain fsync(2) on Darwin does NOT flush the drive’s write cache to media, which would have silently undermined the durability claim on macOS deployments. SQLite, Postgres, and git all use F_FULLFSYNC for the same reason. Skipped on Windows (non-production target). NFS / fuse / network FS no longer claimed as covered — docstring downgraded to “use local FS for production durability claims” (review-driven finding: prior NFSv4 claim overstated; mount options + server semantics make the guarantee non-portable).

Added

  • Retry policies with exponential backoff (sprint 0.3-S5). RetryPolicy { max_attempts, on_transient } on a step is now honored properly: the runner loops up to max_attempts total attempts with 100ms × 2^N (capped at 5s) backoff between. Both sequential and concurrent execution paths share a single retry_with_backoff helper, so retry semantics don’t drift between paths. Final-attempt failure surfaces as "failed after N attempts: <reason>" for operator triage.
  • New library API: boruna_orchestrator::workflow::retry_with_backoff and retry_backoff_ms (pub(crate); used by tests).
  • Operators see retry attempts logged to stderr (gated under cfg(not(test)) so the unit suite stays silent).

Fixed

  • Retry semantics no longer cap at “retry once.” Prior code (should_retry = ... && r.max_attempts > 1) re-attempted exactly once regardless of the configured max_attempts. Now honored as documented: a max_attempts: 5 policy retries up to 4 times.

  • retry_with_backoff’s eprintln gated under cfg(not(test)) (review-driven 0.3-S5 finding #1). Prior unconditional eprintln polluted unit-test stderr and any embedder capturing process stderr.

  • Integration test tests/retry_timing.rs locks real wall-clock backoff (review-driven 0.3-S5 finding #2). Unit tests skip sleeps under cfg(test); this integration test runs in a context where cfg(test) is NOT set on the orchestrator lib build, so the real sleeps fire and the test asserts elapsed >= 250ms for a 3-attempt retry. Catches future regressions that accidentally remove the sleep.

  • Concurrent step execution within a workflow run (sprint 0.3-S4). New --concurrency <N> flag on boruna workflow run and boruna workflow resume. Default 1 = sequential (preserves prior behavior); higher values parallelize fan-out workflows. The per-step output_hash is bit-identical across concurrency levels for successful runs — the determinism contract holds. Locked by a regression test that runs the same workflow at concurrency=1 and concurrency=4 and asserts every step’s hash matches.

  • Implementation: wave-based scheduler. WorkflowValidator::topological_levels partitions the DAG into “waves” where each level’s steps have all dependencies in earlier levels. Within a wave, source steps fan out to short-lived std::thread::spawn’d workers (no tokio, no async runtime). Workers are pure compile+run paths returning a Value; the coordinator owns all DataStore + SQLite mutation.

  • New library API: RunOptions::concurrency: usize, ResumeOptions::concurrency: usize, WorkflowValidator::topological_levels. RunOptions::default() and ResumeOptions::default() initialize concurrency to 1.

  • Persistent path only — WorkflowRunner::run (ephemeral) stays single-threaded. The CLI rejects --concurrency 0 at parse.

Fixed

  • Concurrent chunk halt no longer detaches sibling workers (review-driven 0.3-S4 finding #1). Prior code used ? inside the join loop, which dropped subsequent JoinHandles and detached their threads — those workers continued executing the workflow_dir even after run_persistent returned. Now the join loop collects all JoinHandle::join() results into a Vec before processing, guaranteeing no thread is left running once the function returns.

  • Pre-validate all chunk inputs before marking any Running (review-driven 0.3-S4 finding #2). Prior code interleaved input validation with mark_step_running_clearing_output, so an input failure mid-chunk left earlier siblings Running on disk forever and the next resume re-executed them silently. Now a two-pass structure: pass 1 validates every chunk member’s inputs (no side effects); pass 2 marks all Running atomically and dispatches.

  • Worker panics now produce attributed Failed checkpoints (review-driven 0.3-S4 finding #3). Prior panic handler only matched &'static str payloads (so panic!("step {} bad", id) fell through to a generic message) and lost the step_id, leaving the panicked step at status=Running on disk. Now: tries String payloads first, carries the step_id alongside each JoinHandle, and records a Failed checkpoint with the panic message.

  • boruna workflow show <run-id> CLI (sprint 0.3-S3). Operator inspection of a single run’s full state: row, step checkpoints with truncated output preview, and approval sentinels. Plain-mode tabular output mirrors workflow list’s aesthetic; --json emits a stable pipe-friendly document for jq consumers. Returns RunNotFound for unknown ids (project-conventions §1).

  • New library API: boruna_orchestrator::workflow::{show_run, RunDetail, ApprovalView}. RunDetail carries a metadata_parse_error: Option<String> field so corrupt-metadata signals reach pipeline consumers (review-driven 0.3-S3 H5: stderr warnings are silently dropped when stdout is piped).

Fixed

  • Atomic-rename in DataStore::store_output (sprint 0.3-S3, closes H4 deferral from 0.3-S2c). Replaces the previous std::fs::write (non-atomic) with tempfile::NamedTempFile::persist. Concurrent readers — including another resumed run process — see either the old contents or the new contents, never a partial torn write. Process-crash safe; full power-loss safety still requires a parent-directory fsync, documented honestly in the method docstring as the next hardening pass.

  • output_hash now equals sha256sum result.json (review-driven 0.3-S3 H2/H3). Previously hash_value used compact JSON while store_output wrote pretty-printed JSON, so an operator running sha256sum runs/<id>/outputs/<step>/result.json got a different hex than the persisted output_hash column — a UX footgun. All three (the hash input, the on-disk file bytes, and the step_checkpoints.output_json SQL column) are now the same compact serialization. Locked by a regression test that compares sha256sum-equivalent of the on-disk bytes against hash_value.

  • workflow show --json no longer panics on multi-byte UTF-8 in step output (review-driven 0.3-S3 C1). Prior code did &output_json[..200] to truncate the preview field, which panicked if byte index 200 landed inside a multi-byte codepoint. New truncate_at_char_boundary helper snaps to the nearest char boundary at-or-below the byte budget. Locked by 4 regression tests covering pure ASCII, exact-boundary, multi-byte-at-boundary, and pure-multi-byte content.

  • Approval-gate completion CLI (sprint 0.3-S2c). Three new boruna workflow subcommands close the operator UX deferred from 0.3-S2b:

    • boruna workflow approve <run-id> <step-id> --data-dir <PATH> — records an approval sentinel in the run’s metadata.approvals.<step>.
    • boruna workflow reject <run-id> <step-id> [--reason <STR>] --data-dir <PATH> — records a rejection sentinel; the optional reason surfaces as the step’s error_msg on resume.
    • boruna workflow list [--status <STATUS>] [--json] --data-dir <PATH> — lists runs ordered by (workflow_name, run_id), optionally filtered by running / paused / completed / failed. After approve, the operator runs boruna workflow resume <run-id> to advance the gate to Completed (with a synthetic empty-record output whose hash is locked by a regression test) and execute downstream steps. After reject, resume halts the run as Failed with the recorded reason.
  • Approval sentinel mechanism on metadata.approvals. The runner’s PersistedRunMetadata now carries a BTreeMap<step_id, ApprovalDecision>. Each decision records decision (approved/rejected), decided_at_ms (operational only — does not feed any audit hash), and an optional human-readable reason. Backward compatible with 0.3-S2b databases: the field defaults to empty if absent.

  • New library API: boruna_orchestrator::workflow::record_approval_decision, list_runs, ApprovalKind, plus error variants StepNotFound, StepNotAtApprovalGate { current_status }, StepAlreadyDecided { prior_decision }, NotAnApprovalGateStep, RunNotResumable { terminal_status } (project-conventions §1).

  • New boruna_orchestrator::persistence::{get_run_metadata, update_run_metadata, compare_and_swap_metadata, list_runs} methods. compare_and_swap_metadata is the atomicity primitive for the approve/reject flow’s read-validate-write cycle.

Fixed

  • Race in record_approval_decision (review-driven, 0.3-S2c). Previous implementation’s read+validate+write spanned three separate SQL transactions; two concurrent operators could both pass the in-memory prior-decision check and silently overwrite each other’s decision. Now wrapped in a CAS retry loop via the new compare_and_swap_metadata primitive — exactly one writer succeeds; the others surface a clean StepAlreadyDecided error. Locked by a 4-thread regression test asserting “exactly 1 ok, 3 already-decided.”

  • Resume halt-cause attribution. When both an independently-failed step (e.g. from a crashed prior run) and a rejected approval gate exist for the same run, the resume’s halt_with_failed_step now preserves the FIRST failure (the actual root cause the operator should chase) rather than overwriting with the gate rejection.

  • Sentinel for non-awaiting_approval checkpoint now emits an explicit eprintln! warning rather than silently no-op’ing, so operators see when their approval doesn’t apply (e.g., pre-approval for a step the workflow hasn’t reached, or stale sentinel for an already-terminal step).

  • Defense-in-depth StepKind::ApprovalGate re-validation in resume. Synthetic empty-record output is now refused for non-gate steps even if a sentinel slipped past record_approval_decision’s validation (e.g. via a future code path bypass). Surfaces as WorkflowRunError::Internal.

  • Persistent workflow runs survive process restarts (sprint 0.3-S2b). Wires the SQLite-backed RunCheckpointStore shipped in 0.3-S2a into WorkflowRunner. New boruna workflow run --data-dir <PATH> writes a runs.db and a checkpoint at every step transition. New boruna workflow resume <run-id> picks up where a crashed or paused run left off — already-Completed steps are restored from persisted output; Running-status checkpoints (mid-step crashes) are re-executed since the runner trusts only Completed. Failed step checkpoints in a non-terminal run halt the resume rather than silently advancing past them (review-driven regression). New --ephemeral flag opts out of persistence; --data-dir falls back to $BORUNA_DATA_DIR then ./.boruna/data. Refuses to resume against a workflow whose hash has drifted (error_kind: workflow_hash_mismatch) and against a missing run_id (run_not_found). The boruna workflow approve CLI shipping in 0.3-S2c will let operators advance approval gates; until then a paused approval-gate run resumes by re-pausing.

  • Deterministic run_id derivation (project-conventions §16). Replaces the wall-clock-keyed format!("run-{name}-{utc now}") with sha256(workflow_hash || ":" || inputs_hash || ":" || counter)[..16] hex. The counter is COUNT(*) FROM runs WHERE workflow_hash = ? read inside an explicit BEGIN IMMEDIATE transaction (review-driven from the initial unchecked_transaction DEFERRED-default race) so concurrent writers either see distinct counter values or hit BUSY and retry. Locked by a multi-thread regression test that fans out 8 concurrent insert_run_with_derived_id calls and asserts all 8 produce distinct ids. Algorithm locked by a golden-vector test computed externally.

  • RunRecord and RunOperational view structs on RunCheckpointStore. Replay-verified columns vs. operational metadata are now structurally distinct types: audit/replay code paths consume RunRecord (no started_at, no updated_at, terminal-only Option<RunStatus>); status dashboards consume RunOperational. Closes the H1 review finding from 0.3-S2a. The original RunRow is retained for back-compat callers.

  • New WorkflowRunner API: run_persistent(def, options, data_dir), resume(run_id, data_dir, options), and ResumeOptions { policy, record, live, workflow_dir_override }. ResumeOptions::policy = None defaults to the persisted policy from the original run (review-driven H2 fix; without this default the CLI’s --policy omission silently collapsed to deny-all).

  • New boruna-cli feature flag persist-sqlite (on by default) that forwards to boruna-orchestrator/persist-sqlite. CLI surfaces a typed error rather than silently downgrading when the flag is off and a persistent run is requested (project-conventions §1).

Fixed

  • Reject-at-parse footgun on persistent runs without the SQLite feature. Previously, cargo build --no-default-features produced a CLI that silently ran boruna workflow run dir --data-dir /tmp/x ephemerally, creating no runs.db and giving the operator no signal. Now the CLI errors with a clear “rebuild with default features, or pass --ephemeral” message.

Added

  • Versioned capability identity (#3, sprint 0.3-S11). New boruna capability list [--json] CLI subcommand and boruna_capability_list MCP tool report a stable capability_set_hash over the binary’s capability surface. Integrators use it as part of a cache key — (source_hash, policy_hash, capability_set_hash, policy.schema_version) — to safely memoize deterministic run results across binary upgrades. Algorithm, caching recipe, and per-capability versioning rules documented in docs/reference/capability-identity.md. All 10 shipped capabilities start at contract version "1".
  • New library API in boruna-bytecode: Capability::ALL (canonical sorted iteration), Capability::version(), CapabilityIdentity, CapabilitySetReport, compute_capability_set_hash(), capability_set_report().
  • protocol_version: 1 field on every boruna-mcp tool response (#6, sprint 0.5-S4, pulled forward from 0.5.0 because FleetQ blocked on it for their validate-on-save UX). Wire-format version of the response envelope; bumps only on breaking shape changes (additive changes keep the version). Locked by crates/boruna-mcp/src/tools/mod.rs::TOOL_RESPONSE_PROTOCOL_VERSION and a 16-case regression test asserting every tool’s success and failure path carries it. Versioning policy and bump rules documented in docs/reference/mcp-server.md under “Stability”. Pairs with Policy.schema_version shipped in 0.2.0.
  • MCP Server Tool Reference documentation at docs/reference/mcp-server.md — wire contract for all 10 boruna-mcp tools: parameter names and types, return shapes, error_kind values, encoding rules, and limits. Driven by FleetQ implementer feedback (post-v0.2.0 follow-up): integrators previously had to read crates/boruna-mcp/src/server.rs to learn that boruna_run’s parameter is source (not script) and that there is no input parameter. Linked from docs/README.md.
  • Structured resource limits in boruna_run (#5, sprint 0.3-S10, FleetQ P1). New optional limits parameter on the MCP boruna_run tool accepting max_wall_ms, max_output_bytes, and max_memory_mb. Overruns return a typed error_kind: "limit_exceeded" with a limit_kind discriminator ("wall_ms" or "output_bytes"), the configured limit, and a human-readable message — so callers can surface clean per-limit UX instead of parsing error strings. max_memory_mb is accepted in the schema but not enforced in 0.3.x (documented as platform-best-effort pending Linux setrlimit work in a future sprint).
  • New boruna-vm::error::VmError::WallTimeExceeded(u64) variant and Vm::set_max_wall_ms(Option<u64>) setter. Wall-clock checked every 1024 steps inside the execute loop; uses std::time::Instant (not chrono::Utc::now() per ADR 001 determinism contract). Wall-time enforcement is wall-clock-keyed and therefore non-deterministic on overrun by construction — max_steps remains the deterministic ceiling; max_wall_ms is the operational guardrail.
  • Output JSON Schema validation gate in boruna_run (#8, sprint 0.5-S6, pulled forward from 0.5.0 because FleetQ wanted it in their pipeline). New optional output_schema parameter on the MCP boruna_run tool accepting any JSON Schema 2020-12 object. The script’s result is validated post-execution; mismatches return error_kind: "validation_failed", phase: "output_validation" with per-path JSON Pointer errors. Malformed or oversized schemas (>256 KB) return error_kind: "invalid_output_schema". Schemas declaring a non-2020-12 $schema are rejected (same “reject at parse, don’t silently override” pattern as 0.3-S10’s unsupported_limit). Error array capped at 100 entries with truncated and total_errors fields. Known limitation: records/enums emit as wrapper objects; schemas for the natural shape will fail. Best for primitive returns. See docs/design-output-schema.md.
  • New jsonschema = "0.30" dependency in boruna-mcp (default features off — no resolve-http or resolve-file, so $ref to remote URLs cannot trigger SSRF or arbitrary file reads).
  • Record/replay for net.fetch (#7, sprint 0.5-S7, pulled forward from 0.5.0). Boruna scripts are deterministic by design; external HTTP is not. New CLI flags on boruna run:
    • --record-net-to <FILE> (requires --live) makes real HTTP calls and persists each (method, url, request_body) → response_body transaction to a sidecar JSON tape file.
    • --replay-net-from <FILE> serves responses from a loaded tape with no real network access. Strict ordered match on (method, url, request_body); mismatch returns a typed error naming the position and differing field; tape exhaustion returns a typed error; under-consumption is silently OK.
    • Mutually exclusive (clap conflicts_with). If --live is set alongside --replay-net-from, replay wins (no real calls happen).
  • New module boruna_vm::net_record_replay (feature-gated under http) exposing NetTransaction, NetTape, RecordingHttpHandler, ReplayingHttpHandler, and TAPE_FORMAT_VERSION.
  • RecordingHttpHandler::with_save_path() arms save-on-drop; the CLI also probes write access on the tape path before the run starts so a CI pipeline like record-net-to fixtures/x.tape && verify x.tape fails fast on disk errors instead of silently producing a stale fixture (review-driven hardening).
  • New shared parser boruna_vm::http_handler::parse_net_fetch_args() used by both the real handler and the recording layer so they can’t silently drift in arg interpretation.
  • Documentation: docs/design-net-record-replay.md (tape format, match strategy, CLI surface, known limitations).
  • Per-call OpenTelemetry observability (#9, sprint 0.4-S5, the LAST FleetQ ask). Always-on tracing instrumentation in CapabilityGateway::call emits boruna.cap spans with attributes cap.name, bytes_in, bytes_out, cap.budget_remaining, error.kind (set on the failure path: denied / budget_exceeded / runtime_error). When no subscriber is installed (the default), span macros are essentially no-ops — zero runtime cost.
  • telemetry Cargo feature on boruna-vm (and mirror feature on boruna-cli) adds an OpenTelemetry OTLP-over-HTTP exporter (opentelemetry 0.27 + opentelemetry-otlp 0.27 + tracing-opentelemetry 0.28). New helper boruna_vm::init_telemetry() reads OTEL_EXPORTER_OTLP_ENDPOINT (and optional OTEL_SERVICE_NAME, defaulting to "boruna"); returns a Disabled no-op handle when the endpoint is unset (Boruna behaves identically to a non-telemetry build), installs the exporter when set. Returns a TelemetryHandle whose Drop flushes pending spans.
  • CLI integration: boruna-cli built with --features telemetry starts a tokio runtime in main, calls init_telemetry() BEFORE parsing CLI args, holds the handle for the binary lifetime, and on shutdown drops the handle THEN drains the runtime with a 5-second timeout (so in-flight OTel HTTP POSTs complete instead of being killed by process::exit).
  • New documentation: docs/design-otel.md (span shape, attribute table, determinism contract, library-version pin set, BYO-subscriber fallback path).
  • boruna_orchestrator::persistence::RunCheckpointStore — SQLite-backed workflow checkpoint store (sprint 0.3-S2a). Implements ADR 001 step 1–5: schema, Connection setup with mandatory PRAGMAs (journal_mode=WAL, synchronous=NORMAL, foreign_keys=ON, busy_timeout=5000), CRUD operations (insert_run, update_run_status, get_run, list_runs_by_status, upsert_step_checkpoint, list_step_checkpoints), and a BEGIN IMMEDIATE retry policy that handles both SQLITE_BUSY and SQLITE_LOCKED with exponential backoff (10ms→50ms→250ms→1.25s) before failing with PersistenceError::Busy. Not yet wired into WorkflowRunner — that integration lands in 0.3-S2b (along with boruna workflow resume <run-id> and --data-dir).
  • New persist-sqlite Cargo feature on boruna-orchestrator (default-on). Adds rusqlite = "0.32" with the bundled feature so SQLite compiles from C source — preserves the musl-static-binary story per ADR 001.
  • Schema embedded via include_str!("schema_v1.sql"). Single-row schema_version table with CHECK (id = 1) constraint structurally prevents stale-row accumulation across migration attempts.
  • PersistenceError::NotFound { entity, key } returned by update_run_status when the target run_id does not exist (review- driven; silent-no-op was rejected as a footgun for the resume path).
  • upsert_step_checkpoint uses COALESCE(excluded.X, existing.X) for started_at, output_json, output_hash so a partial upsert (e.g. step transition from Running to Completed without re-supplying started_at) preserves the original value rather than clobbering to NULL (review-driven; locked by two regression tests).
  • docs/design-persistence-store.md — sprint scope split rationale, acceptance criteria, schema annotation conventions.

Decided

  • ADR 001 — Persistence Backend (docs/adr/001-persistence-backend.md). SQLite via rusqlite/bundled chosen as the workflow-checkpoint backend. No persistence-trait abstraction in v1 — direct concrete dependency. Includes a determinism contract for persisted state (operational vs. replay-verified columns), the writer serialization model, mandatory connection PRAGMAs (journal_mode=WAL, foreign_keys=ON, busy_timeout=5000), and an illustrative schema. Unblocks 0.3-S2 through 0.3-S9 — the entire 0.3.0 critical path. Sprint 0.3-S1.

0.2.0 - 2026-04-25

Driven by implementer feedback from FleetQ (production integrator). This release closes the two P0 adoption blockers; remaining P1/P2 asks are tracked as issues #3–#9.

Added

  • MCP boruna_run tool now accepts a structured Policy object for the policy parameter, in addition to the existing "allow-all" / "deny-all" string shorthands. This exposes the per-capability rules (allow, budget), default_allow mode (allowlist vs. denylist), and net_policy (allowed domains, methods, byte limits, timeout) that the VM has always supported. See docs/reference/policy-schema.md and docs/reference/policy.schema.json.
  • New documentation: docs/reference/policy-schema.md (prose + examples) and docs/reference/policy.schema.json (machine-readable JSON Schema 2020-12) for integrators rendering capability matrices in their own UIs.
  • The boruna_run MCP tool description now advertises the structured-policy capability so AI agents discover it from the tool list directly.
  • Multi-target release workflow (.github/workflows/release.yml) that publishes static binaries on every v* tag for x86_64-unknown-linux-musl, aarch64-unknown-linux-musl, x86_64-apple-darwin, and aarch64-apple-darwin, plus a combined SHA256SUMS checksum file. Linux builds use musl so the binaries run on Alpine and other libc-minimal distributions.
  • docs/releasing.md — release process, verification, and rationale for using GitHub-hosted runners (vs. the self-hosted runner used by ci.yml).
  • README install section showing curl-and-verify install.

Changed

  • Breaking (MCP only): boruna_run now rejects unknown policy values (e.g. typo’d strings, numbers, arrays) with success: false, error_kind: "invalid_policy" instead of silently treating them as "allow-all". The legacy strings "allow-all" and "deny-all" continue to behave identically.

0.1.0 - 2026-02-21

Added

  • Deterministic workflow execution engine with DAG validation and topological ordering
  • Hash-chained audit logs (SHA-256) and self-contained evidence bundles for compliance
  • Policy-gated capability system — 10 capabilities: net.fetch, db.query, fs.read, fs.write, time.now, random, ui.render, llm.call, actor.spawn, actor.send
  • Replay engine for determinism verification via EventLog comparison
  • Three reference workflow examples:
    • llm_code_review — linear 3-step pipeline demonstrating LLM capability and evidence recording
    • document_processing — fan-out/merge 5-step pipeline demonstrating parallel steps and DAG scheduling
    • customer_support_triage — approval-gate 4-step pipeline demonstrating human-in-the-loop and conditional pause
  • MCP server (boruna-mcp) exposing 10 tools over JSON-RPC stdio for AI coding agent integration
  • Actor system with OneForOne supervision and bounded execution scheduling (Vm::execute_bounded)
  • boruna-tooling: diagnostics with source spans, auto-repair, trace-to-tests, stdlib test runner, 5 app templates
  • boruna-pkg: deterministic package system with SHA-256 content hashing, dependency resolution, and lockfiles
  • Real HTTP handler (feature-gated via boruna-vm/http) with SSRF protection for net.fetch capability
  • CLI binary (boruna) with subcommands: compile, run, trace, replay, inspect, ast, workflow, evidence, framework, lang, trace2tests, template
  • Standard library: 11 deterministic libraries — std-ui, std-forms, std-authz, std-http, std-db, std-sync, std-validation, std-routing, std-storage, std-notifications, std-testing
  • 557+ tests across 9 crates