diff --git a/README.md b/README.md index 2f0bf5b..e46049d 100644 --- a/README.md +++ b/README.md @@ -68,10 +68,12 @@ The operations cover read-write Linked Data, SPARQL queries, URI manipulation, a - `Value` - `Variable` - `ForEach` + - `Iterate` - `Filter` - `Bindings` - `Current` - - `Execute` + - `Position` + - `Last` - `Merge` - LinkedDataHub-specific - `ldh-CreateContainer` @@ -81,6 +83,7 @@ The operations cover read-write Linked Data, SPARQL queries, URI manipulation, a - `ldh-AddGenericService` - `ldh-AddResultSetChart` - `ldh-AddSelect` + - `ldh-AddConstruct` - `ldh-AddView` - `ldh-AddObjectBlock` - `ldh-AddXHTMLBlock` diff --git a/architecture-evaluation.md b/architecture-evaluation.md new file mode 100644 index 0000000..0aa0092 --- /dev/null +++ b/architecture-evaluation.md @@ -0,0 +1,390 @@ +# Web-Algebra Architecture Evaluation + +*Scope: Web-Algebra (this repo, v1.5.0, Python/rdflib) evaluated on the conceptual and +code level, and compared against its Java/XML sibling `../REST-VKG` (webalgebra module +v1.12.0-SNAPSHOT, Jena/Saxon). Written 2026-07-12.* + +## Verdict in brief + +The premise — an LLM compiling a whole Linked Data workflow into a declarative, +composable operation document instead of issuing step-by-step tool calls — is sound, +and was ahead of its time. The weak points are not the idea but its contracts: +`formal-semantics.md` is a type *catalog* rather than a semantics, the two sibling +implementations have quietly forked (both in operations and in execution model), and +the Python codebase has a handful of structural debts (mutable shared state, a +god-class `Operation`, a three-way execution surface) that are cheap to fix now and +expensive later. The rdflib-based data model is the right choice and its JSON-LD +boundary discipline is genuinely well designed. Part II proposes a three-tier plan: +unify the spec, sync the operation set both ways, then pay down the Python debts. + +--- + +# Part I — Evaluation + +## 1. The premise: composed operations as LLM-emitted bytecode + +**Sound? Yes.** The core bet (README: agents "compile entire workflows into optimized +JSON 'bytecode' that executes atomically") separates *planning* from *execution*: + +- One LLM turn produces the whole program; execution is then deterministic, cheap, and + free of per-step model round-trips. For N-row ForEach workloads this is the + difference between 1 LLM call and O(N) calls. +- The artifact is inspectable and replayable — a JSON document you can review, diff, + version, and re-run, which per-step tool calling can never give you. +- The value domain is RDF-native (`URIRef`/`Literal`/`Graph`/`Result`), so data flows + between operations *as RDF*, not as strings squeezed through generic tool-call JSON. + This is the part most "agent + SPARQL endpoint" designs get wrong. + +**Innovative? Yes, with context.** The now-mainstream pattern of "have the model emit +a program over the tool surface instead of chaining tool calls" (CodeAct-style agents, +code-mode MCP execution) arrived after this design. Web-Algebra's distinctive +contributions beyond that pattern are: + +1. **A domain algebra, not a general-purpose language.** The constrained operation set + is what makes documents verifiable, replayable, and safe — an LLM emitting Python + can do anything; an LLM emitting Web-Algebra can only do what the algebra permits. +2. **XSLT lineage for control flow** — `ForEach`/`Value`/`Current`/`Variable` map to + `xsl:for-each`/`xsl:value-of`/`current()`/`xsl:variable`, a proven declarative + iteration model rather than an invented one. +3. **JSON-LD as both code and data carrier** (§3) — the same document embeds RDF + payloads and operations, discriminated by `@`-keys. + +**Where the premise is under-delivered** (these are contract gaps, not design flaws): + +- **No validation phase.** The "bytecode" is never typechecked before running. A type + error in operation 7 surfaces after operations 1–6 have already POSTed/PUT to live + servers; there is no dry-run, and HTTP effects are not transactional. A compiler + metaphor implies a checker; the checker is missing. +- **Quotation is implicit.** `process_json` evaluates nested `@op` eagerly, + depth-first — except `ForEach`, which receives its `operation` argument raw and + evaluates it once per row (`for_each.py:53`). That makes `ForEach` a special form in + the Lisp sense, but nothing in the spec or the schema declares which arguments are + evaluated and which are quoted. Any future conditional/short-circuit op will hit the + same wall. +- **Effects are typed as pure functions.** `POST : URI × Graph → Result` reads like + arithmetic; nothing distinguishes effectful operations, their ordering guarantees + (currently: top-level JSON array order + ForEach row order), or their failure + behavior. + +## 2. The formal definitions (`formal-semantics.md`) + +**What it does well.** Every operation gets a dual signature (abstract + +concrete Python), the abstract type language is compact (`Term = URI + Literal + +BNode`, `Maybe`, `Sequence α`), and the Strict Type Checking property is stated +explicitly. As a catalog it is mostly accurate against the code. + +**What it is not: a semantics.** There are no evaluation rules. The document never +defines: + +- what nesting means (evaluation order, eager vs quoted arguments); +- what a top-level JSON array means (sequencing? variable-stack accumulation? — + `operation.py:130-137` implements accumulation, the spec is silent); +- how ForEach context propagates and how `Value` resolves the name (`$var` stack vs + context lookup is stated informally at best); +- what happens on error (class, propagation, partial results). + +The type system also gives up exactly where the DSL is most interesting: +`Context = Any` (`formal-semantics.md:20`), `VariableStack = [Dict[String, Any]]`, +and the control-flow ops are typed `Any → Any`. The algebra's *data* operations are +precisely typed; its *composition* operations are untyped. + +Concrete defects, all cheap to fix: + +- `Variable : String × Any × VariableStack → ⊥` (`formal-semantics.md:69`) — `⊥` means + non-termination; the intended type is `Unit`. +- Filter's sequence case reads `Sequence α × Expression → α` (`formal-semantics.md:99`) + — should be `→ Sequence α`. +- `Concat` and `ExtractOntology` are implemented but absent from the catalog. +- `tests/SPEC_GAPS.md` tracks **30+ places where the spec underdetermines the + implementation** — datatype of `Str` results, `Substitute`'s variable syntax and + term-serialization rules, ForEach output shape for `None`/sequence inner results, + the entire error-semantics column, and the JSON arg-key names for ~20 operations. + The spec-driven test suite is the best evidence of the gap: dozens of tests are + `pytest.skip("UNCLEAR(spec)")`. + +**The deeper problem: the spec has forked.** REST-VKG carries its own +`docs/WEB-ALGEBRA.md` — 994 lines of *operational* semantics (execution model, +patterns, examples) versus this repo's 360-line type catalog. Neither references the +other; each has drifted toward its host implementation. For two projects that "should +be in sync in terms of the operations and their signatures", the single highest-value +move is one shared spec (Part II, Tier 1). + +## 3. The JSON DSL + +**Strengths.** + +- `@op`/`args` rides JSON-LD's established `@`-keyword convention, so one document + carries both operations and RDF data with unambiguous discrimination. +- The best design element in the codebase is the **quasi-quotation of JSON-LD + bodies**: a dict carrying `@context`/`@graph`/`@id`/`@type` is treated as data, but + `_resolve_jsonld` (`operation.py:144-175`) walks it and evaluates embedded `@op` + holes in place — e.g. an `@id` computed from a runtime binding — while deliberately + *not* parsing to a `Graph`, so the consuming operation can parse with its own base + IRI. The reasoning is documented in the code (`operation.py:102-119`) and it is + correct: central parsing would freeze unresolved holes into blank nodes. +- Flat, schema-describable JSON is arguably *easier* for an LLM to emit correctly than + free-form code, and trivially validatable — once a validator exists. + +**Weaknesses.** + +- **Verbosity.** Every URI-valued argument needs a nested + `{"@op": "URI", "args": {"input": ...}}` cast, every string interpolation a + `Concat`/`Value` tree. In `examples/united-kingdom-cities.json`, a two-step workflow + costs 104 lines and 14 operation nodes, roughly a third of which are casts and + variable plumbing. This is a *type-system choice* (§5) — plain JSON strings become + `xsd:string` literals (`operation.py:275-277`), so URIs must be cast explicitly — + but the cost lands on every document and every LLM emission. Sugar is available + without weakening the typing: `{"@id": "..."}` is already valid JSON-LD for "this is + a URI" and could be accepted anywhere a URI is expected. +- **No envelope.** Documents have no version, no namespace, no name — just a bare + array. The Java XML side has `xmlns="https://w3id.org/atomgraph/web-algebra"`; the + JSON side has nothing to dispatch or validate against. +- **The JSON-LD sniff** (`_JSONLD_KEYS`, `operation.py:17`) is a heuristic: any dict + containing `@type` is data. The code comment argues these keys are unambiguous, and + within RDF workflows that mostly holds, but it is a global reserved-word rule the + spec never states. +- **Arrays are overloaded**: top-level array = sequential program with shared variable + stack; array under `ForEach.operation` = per-row sequence where only the last + non-`None` result is kept (`for_each.py:84-97`); array elsewhere = plain argument + list. Three meanings, zero spec lines. +- **JSON arg keys are folklore.** The spec gives Python parameter names; the JSON + layer uses different keys (`select` vs `select_data`), confirmed only by fixtures + (SPEC_GAPS "JSON dispatch surface"). + +## 4. The Python codebase + +Ranked by how much they matter: + +1. **Mutable default arguments** — `execute_json(self, arguments, variable_stack: + list = [])` (`operation.py:56`), `process_json(..., context: dict = {}, + variable_stack: list = [])` (`operation.py:83-84`), and the class field + `context: Any = {}` (`operation.py:31`). Python evaluates these once; every call + that omits the argument shares one list/dict across the process. Today the + single-threaded interpreter mostly masks it; the day two documents run in one + process (the MCP server is exactly that), variables leak between executions. This + is the one outright *bug class* in the core. +2. **`Operation` is a god class.** Registry, the whole interpreter (`process_json` + + `_resolve_jsonld`), variable-stack management, and five type-conversion helpers all + live in the ABC every operation inherits (`operation.py`, 345 lines). The + interpreter is not an operation concern; it should be a separate module holding an + execution context — which is also the precondition for parallelism (§6). +3. **The `execute()` contract is violated by its own flagship op.** + `ForEach.execute()` raises `NotImplementedError` (`for_each.py:42-44`) because the + pure layer has no way to run an operation with context — i.e. the algebra's central + higher-order operation has no pure form, only a JSON form. `Bindings` and + `ldh-List` return `list[dict]`, many `ldh-*` ops return `Any`. The abstract + signature `execute(*args)` is variadic while every implementation is fixed-arity, + so type checkers verify nothing. +4. **pydantic is decorative.** `extra="allow"`, no field validation, hand-written + `inputSchema()` dicts that often omit `"type"` (`for_each.py:24-34`) and are never + used to validate anything. The DSL's missing validator (§1) could be generated from + real pydantic models; today neither exists. +5. **Three execution surfaces per operation** (`execute` / `execute_json` / + `mcp_run`) is a triple maintenance burden, and it shows: `ForEach.mcp_run` returns + the static string "ForEach operation completed" (`for_each.py:112-114`). 21 of 43 + ops are MCP-exposed; the boundary between "MCP tool" and "DSL-only" is undocumented. +6. **No exception taxonomy.** `TypeError` vs `ValueError` vs raw + `urllib.error.HTTPError` varies by op; SPEC_GAPS' error-semantics section exists + because callers cannot classify failures. +7. **Duplication.** Seven HTTP-backed ops repeat the same `LinkedDataClient` + construction in `model_post_init`; isinstance-check boilerplate opens every + `execute()`. +8. **Minor:** `_serialize_for_json_context` (`operation.py:177-188`) is dead code; no + retry on transient network failures (only 429 `Retry-After` is honored). + +**What is genuinely good** — worth saying plainly, because it should be preserved +through any refactor: + +- The **JSON-LD → Graph boundary** is exactly right: `to_graph()` + (`operation.py:215-242`) is the single parse point, each op applies its own base + IRI, and the design rationale is written down where it matters. +- `client.py` handles 308 redirects and 429 backoff correctly, with client-cert auth + cleanly isolated in settings. +- The registry + auto-discovery pattern is clean and scales. +- The **spec-driven test discipline** (203 tests authored from the spec alone, + `SPEC_GAPS.md` recording every ambiguity with a proposed spec edit) is a practice + most projects never reach. The gaps it found are the spec's problem, not the suite's. + +## 5. The rdflib data model + +Right choice, well executed at the boundaries. `Node`/`Graph`/`Result` as the value +domain keeps datatype and language-tag fidelity end-to-end; `JSONResult` +(`json_result.py`) is a clean adapter between rdflib results and the SPARQL JSON +wire format; `POST`/`PUT` returning a `Result` of `{status, url}` bindings is a nice +touch — HTTP responses become queryable data, and it happens to match REST-VKG's +`ResultSet` shape exactly. + +Two observations: + +- **Leaf typing is principled and expensive.** Every plain JSON string becomes + `Literal(..., datatype=xsd:string)` — never a URI (`operation.py:275-277`). That + strictness is defensible RDF hygiene (Java's `String`-typed leaves lose + datatype/lang information, see §6) and it is what forces the `URI` cast operation + and much of the DSL's verbosity. Keep the typing; add sugar at the JSON boundary. +- The stated convention "execute() is pure rdflib" holds for the data-plane ops but + not the control plane (§4.3). Either the convention gets a carve-out for + interpreter-level forms (ForEach, Execute, Variable, Value, Current, Filter, + Bindings) — which is honest, they are special forms, not term functions — or those + need pure formulations. The spec should say which. + +## 6. Python ↔ Java: parity and divergence + +### Operation parity (registry names) + +| | Operations | +|---|---| +| **Shared core (20)** | GET, POST, PUT, SELECT, DESCRIBE, CONSTRUCT, Merge, ForEach, Value, Str, Concat, Replace, EncodeForURI, ResolveURI, Substitute, Variable, Current, Execute, STRUUID/StrUUID, SPARQLString | +| **Java-only (1)** | Iterate — stateful pagination: `param` initialization, `next-iteration` passing, `break` conditions (`IterateOperation.java`, 244 lines) | +| **Python-only, generic (9)** | PATCH, Values, URI, Filter, Bindings, ExtractClasses, ExtractDatatypeProperties, ExtractObjectProperties, ExtractOntology | +| **Python-only, product-specific (14)** | the `ldh-*` LinkedDataHub operations | + +(Java's GRDDL is superseded by the client-side response filter and excluded.) + +### Three deep divergences — and per-op verdicts + +**a. ForEach: map → List (Python) vs parallel map-merge → Model (Java).** +Python's ForEach returns the list of per-row results (`for_each.py:79-110`); Java's +requires the select to yield a `ResultSet`, requires the inner operation to return a +`Model`, executes rows via `parallelStream()`, and merges into one `Model` +(`ForEachOperation.java:61-100`). **Verdict: Python's shape is the better algebra; +Java's execution model is the better runtime.** `Sequence α × (α → β) → Sequence β` +is more general — Java's fused map-merge cannot express "PUT each row's document and +give me the statuses" (the UK-cities example) without contortion, and its two +`instanceof` gates are exactly the kind of restriction a spec should not bake in. +Java's merge is `Merge(ForEach(...))` — an explicit composition Python already has. +Conversely, Java's per-row **immutable context clone enabling parallel iteration** is +strictly better than Python's shared mutable stack. Sync direction: spec ForEach as +sequence-returning with the map-merge documented as a fused specialization Java may +keep; Python adopts context isolation (and then parallelism) from Java. + +**b. Variables/context: mutable stack (Python) vs immutable `ExecutionContext` +(Java).** Java's context is a persistent structure — `withVariable`/`withBinding` +return new instances (`ExecutionContext.java:81-98`), with a progress emitter riding +along (`start:`/`complete:`/`error:` events). **Verdict: Java wins outright.** This is +thread safety, ForEach-row isolation, and observability in one move, and it is the +enabler for fixing Python issues §4.1 and §4.2 in a way that converges the two +codebases instead of diverging them further. + +**c. Leaf typing: RDF terms (Python) vs Strings (Java).** Java string ops return +`String` and `Value` stringifies RDF nodes; Python returns typed `Literal`s and keeps +`URIRef`/`Literal` distinct end-to-end. **Verdict: Python wins.** A `String`-typed +data plane silently drops datatypes and language tags — the exact failure RDF systems +exist to avoid. Long-term, Java should move its operation returns toward +`RDFNode`-typed values; the shared spec should define signatures in abstract RDF terms +(as `formal-semantics.md` already does) either way. + +### Maturity gaps (Java ahead, no controversy) + +- **Iterate** — no Python equivalent for paginated APIs; the biggest functional gap. +- **Parallel ForEach** — blocked in Python only by the mutable context. +- **Progress events** — Python has `logging.info` only; no structured lifecycle. +- **SPARQLString hardening** — Java takes question + endpoint, injects the service's + AGENTS.md plus date/timezone into the prompt, and retries 3× with exponential + backoff; Python takes a bare question with no retries. +- **Hybrid SELECT** — Java's SELECT accepts a remote endpoint *or* a local graph; + Python is endpoint-only, which blocks querying intermediate in-memory results. +- Deployment/observability (Docker, health endpoints, timing metrics) — product-level + rather than algebra-level, but worth noting. + +--- + +# Part II — Implementation plan + +Three tiers, independently executable, in value order. Decisions already made: +divergences are resolved per-op as recommended above; parity scope is the generic +core in both directions (`ldh-*` stays Python-only, declared as a product extension); +all three tiers are in scope. + +## Tier 1 — One spec, made whole (highest value, zero code risk) + +1. **Unify the fork.** Merge this repo's `formal-semantics.md` (type catalog) and + REST-VKG's `docs/WEB-ALGEBRA.md` (operational semantics) into a single canonical + spec shared by both repos (one home, the other references it — natural candidate: + a spec file under the `w3id.org/atomgraph/web-algebra` namespace both already + implicitly claim). Structure: type system → **evaluation rules** → operation + catalog → error semantics → serializations (JSON and XML as two concrete syntaxes + of one abstract syntax). +2. **Write the missing evaluation rules** (the §2 list): eager depth-first argument + evaluation; *quoted operands declared per-op* (ForEach.operation, Execute's body); + top-level array sequencing incl. variable-stack accumulation; ForEach context + propagation and `Value` resolution order (`$var` stack lookup vs context binding); + the three meanings of arrays; effect annotation for GET/POST/PUT/PATCH/SELECT/ + CONSTRUCT/DESCRIBE and their ordering guarantees. +3. **Resolve `tests/SPEC_GAPS.md` item by item** — it already contains proposed edits + for nearly every entry; most are one-line decisions (Str result datatype, + EncodeForURI's RFC, Merge set-semantics, Bindings ordering, error classes, + JSON arg-key table). Fix the two catalog defects (`⊥` → `Unit`, + Filter → `Sequence α`) and add the two missing entries (Concat, ExtractOntology). +4. **Declare the extension model**: `ldh-*` as a named product-specific extension + namespace; Iterate added to the core catalog (from Java). + +*Verification:* every resolved item un-skips its `UNCLEAR(spec)` tests; +`uv run pytest -m 'not network and not sparql and not ldh'` green with strictly fewer +skips than today. + +## Tier 2 — Operation & signature sync (generic core, both directions) + +**Python gains:** + +| Item | Notes | +|---|---| +| `Iterate` | Port from `IterateOperation.java` (params, next-iteration, break condition). Spec first (Tier 1), then implement + spec-driven tests. | +| Hybrid `SELECT` | Accept a `Graph` argument alternative to `endpoint` (query local/intermediate results via rdflib). Mirrors `SelectOperation.java`. | +| `SPARQLString` parity | Add endpoint parameter and context injection; add bounded retry with backoff. | +| Parallel ForEach | After Tier 3.2 (immutable context). Row isolation semantics per unified spec. | + +**Java gains** (tracked here, implemented in REST-VKG): `PATCH`, `Values`, `Filter`, +`Bindings`, `URI`, and the four `Extract*` schema ops — signatures taken verbatim from +the unified catalog. + +**Harmonizations:** + +- ForEach per verdict §6a: spec is sequence-returning; Java either generalizes or its + map-merge is documented as a fused `Merge ∘ ForEach` specialization. +- Registry-name alignment: `STRUUID` vs `StrUUID` — recommend `STRUUID` (matches the + SPARQL function name, which is the naming rule the other ops already follow). +- POST/PUT `{status, url}` result shape: already aligned; codify it in the catalog. + +*Verification:* a parity table in the unified spec, asserted by a test in each repo +that diffs its registry against the spec catalog (Python: registry names vs a +committed list; the SPEC_GAPS re-verify note at `tests/SPEC_GAPS.md:23` already asks +for exactly this). + +## Tier 3 — Python code health + +Ordered so each step stands alone: + +1. **Kill mutable defaults.** `variable_stack: list = None` → `if None: []` (or + required-arg) in `execute_json`/`process_json`/all ops; `context` field default via + pydantic `default_factory`. Mechanical, high value. (`operation.py:31,56,83-84` + and every operation's `execute_json`.) +2. **Extract the interpreter.** Move `process_json`, `_resolve_jsonld`, and + variable-stack handling out of `Operation` into an `interpreter.py` with an + immutable `ExecutionContext` (variables + current binding + progress callback) — + deliberately the same shape as `ExecutionContext.java`, converging the two + codebases. Operations keep `execute`/`execute_json`; `self.context` and the stack + parameter are replaced by the context object. +3. **Exception taxonomy.** `WebAlgebraError` base; `UnknownOperationError`, + `OperationTypeError`, `VariableNotFoundError`, `HttpOperationError(status, url)`. + Raise-sites updated; spec's error-semantics section (Tier 1.2) is the contract. +4. **Deduplicate the HTTP plumbing.** One client factory/mixin for the 7 ops that + construct `LinkedDataClient` in `model_post_init`; add transient-failure retry at + the client, not per-op. +5. **Honest contracts.** Delete `_serialize_for_json_context`; declare the + interpreter-level special forms (ForEach, Execute, Variable, Value, Current) as + such instead of pretending at a pure `execute()` (drop the `NotImplementedError` + stub per the unified spec); give ForEach a real `mcp_run` or remove its `MCPTool` + claim; replace hand-written `inputSchema()` dicts with schemas generated from + pydantic argument models — which also yields the missing pre-execution document + validator (§1) nearly for free. + +*Verification:* full suite `uv run pytest` after each step; step 2 additionally +verified by the integration fixtures in `tests/integration/` (behavior-preserving +refactor); step 5's validator gets new negative fixtures (malformed documents rejected +before any HTTP effect). + +## Suggested sequencing + +Tier 1 first (it unblocks skipped tests and is the sync keystone). Tier 3.1 anytime +(it is a bug fix). Tier 3.2 before the parallel-ForEach item of Tier 2. Everything +else is independent. diff --git a/formal-semantics.md b/formal-semantics.md index 485ce09..1955ebe 100644 --- a/formal-semantics.md +++ b/formal-semantics.md @@ -1,360 +1,1076 @@ # Web Algebra Formal Semantics -## Abstract Type System +This document is the normative specification of Web Algebra: its type system, its +JSON serialization, its evaluation semantics, and the catalog of core operations. +The Python implementation in this repository and the test suite under `tests/` +conform to it; where the two disagree, this document wins and the code is wrong. + +The LinkedDataHub-specific `ldh-*` operations are a product extension and are +described in the informative Appendix A. A sibling implementation (REST-VKG, +Java) serializes the same algebra as XML under the namespace +`https://w3id.org/atomgraph/web-algebra`; this document specifies the abstract +algebra and its JSON serialization. + +## 1. Type System + +### 1.1 Abstract types + +``` +URI = URI reference +Literal = literal value with optional datatype IRI or language tag +BNode = blank node +Term = URI + Literal + BNode +Graph = RDF graph (set of triples) +Result = SPARQL SELECT result: a variable list and an ordered sequence + of Bindings. Result values are materialized — they hold their + rows and may be iterated any number of times +Binding = one solution row: a partial mapping from variable names to Terms +Sequence α = ordered list of values of type α; flat, as in XDM — a sequence + is never an item of a sequence +Position = Literal with datatype xsd:integer and value ≥ 1 (XSLT-style + 1-based index; also the name of the focus-accessor operation, + §4.1 — context disambiguates) +Unit = the empty sequence, written (): the value of an operation + executed for its effect +Context = the current iteration item (see §3.5); one of + Binding + Term + Graph + JSON value +Focus = Context × Position × Position — the dynamic context of an + iteration: (item, position, size), exactly XSLT's focus (§3.5) +Environment = stack of variable scopes; each scope maps names to values +Operation = an unevaluated operation form (see quoting, §3.3) +``` + +### 1.2 Concrete Python types -### Primitive Types -``` -URI = Abstract URI reference -Literal = Abstract literal value with optional datatype and language -BNode = Abstract blank node identifier -Graph = Abstract RDF graph -Term = URI + Literal + BNode -``` - -### Collection Types -``` -Sequence α = [α] -- Ordered list of elements -Result = SPARQL SELECT result with variables and bindings -ResultRow = Single binding row from Result -VariableStack = [Dict[String, Any]] -- Stack of variable scopes -Context = Any -- Current execution context (varies by operation) -``` - -### Operation Types -``` -Operation = Abstract operation that can be executed -Expression = Operation + Literal + Integer -- Expressions for filtering -``` - -## Concrete Python Type System - -### RDFLib Types ```python -URI = rdflib.URIRef -Literal = rdflib.Literal -BNode = rdflib.BNode -Graph = rdflib.Graph -Term = Union[rdflib.URIRef, rdflib.Literal, rdflib.BNode] +URI = rdflib.URIRef +Literal = rdflib.Literal +BNode = rdflib.BNode +Term = Union[rdflib.URIRef, rdflib.Literal, rdflib.BNode] +Graph = rdflib.Graph +Result = rdflib.query.Result # typically web_algebra.json_result.JSONResult +Binding = rdflib.query.ResultRow # or Dict[str, Term] via Bindings +Sequence = list +Unit = None # concatenates as the empty sequence +Environment = list[dict[str, Any]] # the "variable stack" ``` -### Collection Types -```python -Sequence = List[Any] -Result = rdflib.query.Result -ResultRow = rdflib.query.ResultRow -VariableStack = List[Dict[str, Any]] -Context = Any -``` - -### Execution Architecture -```python -# Triple execution pattern - all operations implement: -def execute(*args: RDFLib_types) -> RDFLib_type # Pure function with RDFLib terms -def execute_json(arguments: dict, variable_stack: list) -> Any # JSON processing -def mcp_run(arguments: dict, context: Any = None) -> Any # MCP interface -``` +## 2. Document Model (JSON serialization) + +### 2.1 Documents + +A Web Algebra document is a JSON document in one of two shapes: + +1. **A single form** — most commonly an operation call object. +2. **A program** — a JSON array of forms, evaluated in order (§3.2, *sequence + form*). + +### 2.2 Forms + +Every JSON value in (the program of) a document is a **form**. The kind of a +form is decided *syntactically*, by the first matching rule: -## Operation Catalog +| # | Syntax | Form kind | +|---|--------|-----------| +| 1 | object with an `@op` member | **operation call** | +| 2 | object whose only member is `@id` | **URI reference** | +| 3 | object with any of `@context`, `@graph`, `@id`, `@type` | **RDF data** (JSON-LD) | +| 4 | any other object | **generic object** | +| 5 | array | **sequence** | +| 6 | string, number, boolean | **scalar** | +| 7 | null | invalid (`TypeError`) | -### Core System Operations +**Operation call.** `{"@op": Name, "args": {key: form, ...}}`. `Name` must be a +registered operation name (§4, Appendix A); otherwise the document is invalid +(`ValueError`). `args` may be omitted when the operation takes no arguments. +Argument keys are operation-specific and normative (§4); a missing required +argument raises `KeyError`. -**Value** - Access variables and context values -``` -Abstract: String × Context × VariableStack → Any -Python: def execute(self, name: str, context: Any, variable_stack: List[Dict[str, Any]]) -> Any +**URI reference.** `{"@id": form}` — an object with *exactly one* member named +`@id` — evaluates to a `URI`. The inner form may be a string or any form that +evaluates to a Term; the URI is its lexical form. This is deliberately the +JSON-LD node-reference syntax: a bare node reference carries no triples, so +reusing it as the URI form is unambiguous. It replaces the verbose +`{"@op": "URI", "args": {"input": ...}}` cast in the common case: + +```json +{"@op": "GET", "args": {"url": {"@id": "https://dbpedia.org/resource/London"}}} ``` -**Variable** - Set variables in current scope (XSLT-style) -``` -Abstract: String × Any × VariableStack → ⊥ -Python: def execute(self, name: str, value: Any, variable_stack: List[Dict[str, Any]]) -> None -``` +Inside an *RDF data* form, `@id` keeps its JSON-LD meaning and is **not** +subject to this rule (rule 3 wins because the discrimination happens on the +enclosing document, whose other members mark it as data). -**Current** - Return current context item -``` -Abstract: Any → Any +**RDF data.** An object carrying any of the four reserved JSON-LD keys +`@context`, `@graph`, `@id`, `@type` is RDF data (a JSON-LD document or +fragment), not a structure to evaluate. These are JSON-LD reserved terms with +no meaning in non-RDF JSON; this list is normative and closed. Within an RDF +data form, embedded operation-call objects ("holes", e.g. an `@id` computed at +runtime) are evaluated in place and replaced by their results; every other +value is left untouched, and the form as a whole remains a JSON structure. It +is parsed into a `Graph` only by the consuming operation, which supplies the +correct base IRI (§2.3). + +**Generic object.** Evaluated member-wise: each value is evaluated as a form, +keys are preserved. This is how, e.g., SPARQL JSON term objects (§2.4) with +computed values are written. + +**Sequence.** Evaluated element-wise, in order, in a fresh variable scope +(§3.4). The value is the concatenation of the element values (§3.2). At +document top level this is the *program* shape. + +**Scalar.** Coerced to a Term: + +| JSON | Term | +|------|------| +| string | `Literal` with datatype `xsd:string` | +| integer | `Literal` with datatype `xsd:integer` | +| number parsed as floating-point (e.g. `1.5`, `1.0`, `1e3`) | `Literal` with datatype `xsd:double` | +| boolean | `Literal` with datatype `xsd:boolean` | + +A plain string is *always* a string literal, never a URI; URIs are written +with the URI reference form or produced by `URI`/`ResolveURI`. `null` is not a +valid form. + +### 2.3 Base IRI + +RDF data forms may contain relative IRIs. The consuming operation — the one +that turns the JSON-LD into a `Graph` — resolves them against its *target* +URI (e.g. `PUT`'s `url`), by parsing with that URI as base. There is no +document-global base IRI. Fragment references like `#service` therefore +resolve against the document being written, which is the intended semantics. + +### 2.4 SPARQL JSON term form + +Where an operation expects a Term argument (noted in its catalog entry), the +SPARQL 1.1 Query Results JSON term object is accepted: + +```json +{"type": "uri" | "literal" | "bnode", "value": "...", + "datatype": "...IRI...", "xml:lang": "..."} +``` + +`datatype` and `xml:lang` are optional and only meaningful for `literal`. An +unknown `type` raises `ValueError`. This is a generic-object form whose +member values may themselves be computed by nested operation calls. + +## 3. Evaluation Semantics + +Evaluation follows XSLT/XPath, applied to RDF. Values form flat, XDM-style +sequences; iteration establishes XSLT's focus; variables live in XSLT's +scopes; `ForEach` is `xsl:for-each` — unordered, with the `xsl:result-document` +rule for the updates it contains — and `Iterate` is `xsl:iterate`. SPARQL +contributes the data model the sequences hold — terms, graphs, result rows — +and the term-level function library (§4.2), which is itself XPath's. SPARQL's +own algebra (join, union, filter over solution mappings) is not re-created as +operations; it stays inside query strings. + +### 3.1 Values + +Evaluation maps forms to values in the domain + +``` +Value = Term + Graph + Result + Binding + Unit + Sequence Value + + Object + Data + +Object = JSON object whose member values are Values — the result of a + generic-object form (members are evaluated, so its scalars have + been coerced to Terms) +Data = JSON-LD structure whose evaluated hole positions hold Terms and + whose remaining content is raw, uncoerced JSON — the result of an + RDF data form +``` + +Sequences are flat (§1.1): constructing one concatenates, so a sequence-valued +element contributes its items and `Unit` — the empty sequence — contributes +none, exactly as `()` vanishes in an XPath sequence. A `Result` is a value in +its own right, not a sequence; concatenation never dissolves it into its rows +(`Bindings` does that, explicitly). + +Note the two closures differ deliberately: a generic object's members are +forms and evaluate (§3.2), while an RDF data form's non-hole content is data +for a JSON-LD parser and must stay untouched. A `Data` value becomes a +`Graph` only at an operation boundary, parsed with that operation's base IRI +(§2.3); an `Object` value becomes a Term only where an operation's catalog +entry accepts the SPARQL JSON term form (§2.4). + +### 3.2 Evaluation rules + +Evaluation is **eager and depth-first**: when an operation call is evaluated, +its argument forms are evaluated first, in document order, and the operation +is then applied to the resulting values. The exceptions are **quoted +operands** (§3.3). + +- *Operation call*: evaluate non-quoted arguments (document order), apply the + operation, yield its result. +- *URI reference*: evaluate the inner form; yield `URIRef` of its lexical + form. +- *RDF data*: walk the structure; evaluate embedded operation calls in place; + yield the resulting JSON structure. +- *Generic object*: evaluate each member value; yield the object. +- *Sequence*: push a fresh variable scope; evaluate elements in order; pop the + scope; yield the concatenation of the element values (§3.1). Elements are + evaluated for both value and effect — an element that is an effectful + operation call (§3.6) executes even if its value is never consumed, and a + `Variable` leaves no item, as `xsl:variable` leaves none in a sequence + constructor. +- *Scalar*: yield the coerced Term (§2.2). + +### 3.3 Quoted operands + +Some operations receive an *unevaluated* form — they are special forms in the +Lisp sense, and their quoted operands are evaluated under a different regime +(later, repeatedly, or in a different context). Quoted operands are marked +**quoted** in the catalog. They are exactly: + +| Operation | Operand | Evaluation regime | +|-----------|---------|-------------------| +| `ForEach` | `operation` | once per iteration item, under the focus *(item, position, size)* | +| `Iterate` | `operation` | once per loop iteration, in the loop's environment | +| `Iterate` | `next-iteration` member values | after each iteration's body, in the iteration's environment | + +All other arguments of all operations are eagerly evaluated. An operation not +in this table never sees an unevaluated form. + +### 3.4 Variable environment + +The environment is a stack of scopes. + +- **Binding**: `Variable` binds a name to a value in the *current* (innermost) + scope, creating one if the stack is empty. Rebinding a name in the same + scope overwrites it. +- **Lookup**: `Value` with a `$`-prefixed name searches scopes innermost to + outermost; a miss raises `ValueError`. +- **Scope creation**: a fresh scope is pushed for the duration of + (a) every sequence form, and (b) every `ForEach` iteration. Consequently a + variable bound in a program step is visible to *subsequent* steps of the + same program and to forms nested within them, and ceases to exist after the + sequence ends; a variable bound inside a `ForEach` iteration does not leak + into the next iteration. + +The `$` sigil belongs to the reference syntax of `Value`, not to the variable +name: `Variable` binds `name`, `Value` reads `$name`. Because the sigil +decides the lookup domain, variable and context lookups never shadow each +other. + +### 3.5 Focus + +The **focus** is the dynamic context of an iteration — the triple +*(item, position, size)*, exactly XSLT's dynamic-context triple. It is +established *only* by `ForEach`, which evaluates its quoted `operation` once +per item with the focus *(item i, i, n)* where *n* is the number of items; +within any evaluation of the operand, 1 ≤ position ≤ size. A nested `ForEach` +shadows the outer focus for the extent of its own operand. Outside any +iteration there is no focus, and operations that require one raise +`ValueError`. + +The focus accessors, named after their XSLT/XPath counterparts: + +- `Current` yields the focus item itself (XSLT `current()`). +- `Position` yields the item's 1-based position as an `xsd:integer` Literal + (XPath `fn:position()`). +- `Last` yields the iteration size as an `xsd:integer` Literal (XPath + `fn:last()`). +- `Value` with an unprefixed name looks the name up *in* the focus item. + The item shapes are **closed**: a `Binding` (SPARQL row) yields the term + bound to that variable name; a mapping (e.g. a JSON object item) yields + the member value. A miss, or any other item shape, raises `ValueError`. + +### 3.6 Effects and ordering + +Operations are marked in the catalog as **pure** (no observable effect), +**query** (reads external state: HTTP GET, SPARQL query), or **update** +(writes external state: HTTP POST/PUT/PATCH). `SPARQLString` calls an external +LLM service and is additionally **non-deterministic**, as is `STRUUID`. + +Ordering guarantees: + +- Argument evaluation is depth-first in document order (§3.2); an operation's + effects happen after all its arguments' effects. +- Sequence elements evaluate in order: all effects of element *n* happen + before any effect of element *n+1*. +- `ForEach` yields its result sequence in item order. Whether iterations + execute sequentially or concurrently is implementation-defined, so effects + *across* iterations are **unordered** (within one iteration, sequence + ordering applies). The ordered constructs are the sequence form and + `Iterate`, as the operations of one SPARQL Update request are. +- Updates inside a `ForEach` are allowed under the rule of XSLT's + `xsl:result-document`: two iterations updating the same URI are an error + (`ValueError`, §3.7), as two result documents with one `href` are + (XTDE1490). The URI compared is the one the write reports (§4.4). Within one + iteration the sequence form orders the writes, so a document may be created + and then added to. The error is raised when the second write is reported, + so that write has been performed, as the second result document has been + when XTDE1490 is raised. An update made inside a nested `ForEach` counts + for every enclosing iteration it runs in. The rule is stated on targets, not + methods: two appends to one graph would commute as set unions, but the + algebra cannot see how a server applies a write (an LDH container answers a + `POST` by creating a child; a block append reads a position before it + writes), so it does not assume that any two writes commute. +- A query effect inside a `ForEach` that reads a URI another iteration of the + same `ForEach` updates has an implementation-dependent outcome: it may see + the resource before or after that update. This is XSLT's XTRE1500 (reading + a resource that the same transformation writes), and like it the case is + not detected. A program whose result depends on it is not portable; to + read what a write produced, order the read after the write in one + iteration, or after the `ForEach` in the enclosing sequence. +- `Iterate` is strictly sequential by definition: iteration *k+1*'s + parameters are computed from iteration *k*'s environment. + +There are no transactions: if evaluation fails midway, effects already +performed are not rolled back. + +### 3.7 Errors + +Failures raise Python exceptions per this table (normative): + +| Condition | Exception | +|-----------|-----------| +| unknown operation name in `@op` | `ValueError` | +| `null` form | `TypeError` | +| missing required argument key | `KeyError` | +| argument or operand of the wrong type (any layer) | `TypeError` | +| unknown variable in `$name` lookup | `ValueError` | +| focus-item lookup miss, or no focus established | `ValueError` | +| two iterations of one `ForEach` updating the same URI (§3.6) | `ValueError` | +| a write (`POST`, `PUT`, `PATCH`, and the `ldh-*` updates) answered outside 2xx (§4.4) | `ValueError` | +| `SPARQLString` produced no query of the declared shape, or one of its `context` values is empty (§4.3) | `ValueError` | +| `bindings` of a schema operation without a `subject` variable, or with no rows (§4.6) | `ValueError` | +| `Filter` position < 1 or > length | `ValueError` | +| regular-expression errors in `Replace` — invalid pattern or flags, zero-length-matching pattern, invalid replacement (XPath `err:FORX000*`) | `ValueError` | +| unknown `type` in SPARQL JSON term form | `ValueError` | +| blank node where SPARQL syntax forbids it (`Values` data) | `ValueError` | +| non-RDF response to a Linked Data or SPARQL operation — unsupported media type, missing `Content-Type`, or a body that does not parse as the negotiated format (§4.3–4.4) | `ValueError` | +| HTTP/SPARQL transport failure: no response, or a read (`GET`, a SPARQL query) answered outside 2xx | `urllib.error.HTTPError` / `URLError`, unwrapped | + +Type checking is strict: operations validate their inputs and raise `TypeError` +*before* performing any effect. No implicit casting is performed between Term +kinds; the only implicit conversion anywhere is scalar coercion (§2.2) and +string-compatibility (§4.2). + +Interpreter-level failures (unknown operation, invalid form, unresolved +variable, missing focus) are raised as `WebAlgebraError` subclasses that +*also* inherit the built-in class tabulated above, so the contract here holds +while a caller may `except WebAlgebraError` to distinguish an ill-formed +document from an unrelated bug. Per-operation argument-type validation stays +plain `TypeError`. Transport failures are the deliberate exception: they +propagate unwrapped as noted, never wrapped. + +### 3.8 Formal definition + +This section is the definition; §§3.2–3.7 restate it in prose, and on any +disagreement this section wins. + +**Abstract syntax.** Forms `e` (concrete syntax per §2.2): + +``` +e ::= s scalar: string | integer | double | boolean + | {"@id": e} URI reference + | {k₁: e₁, …, kₙ: eₙ} generic object (no reserved keys) + | data RDF data form (JSON-LD object; may contain + operation-call holes) + | [e₁, …, eₙ] sequence + | op(k₁: e₁, …, kₙ: eₙ) operation call, op a catalog name; operands + marked ⟨quoted⟩ in the catalog are taken as + unevaluated forms +``` + +**Semantic domains.** + +``` +v ∈ Value §3.1 +v ⧺ w sequence concatenation, XDM-style: a sequence + operand contributes its items, () none; the + result is flat (§3.1) +ρ ∈ Env = Scope* stack of scopes; Scope = Name ⇀ Value +φ ∈ Focus⊥ = (Value × ℕ⁺ × ℕ⁺) + ⊥ (item, position, size), or absent +σ ∈ World external web state (graphs behind URIs, + endpoint contents); opaque +``` + +Each catalog operation `op` that is not treated by a rule below is a plain +operator with an interpretation + +``` +δ_op : Value* × World → (Value × World) + Err +``` + +*Pure* operations neither read nor write the World; *query* operations read +it; *update* operations read and write it; `STRUUID` and `SPARQLString` are +relations rather than functions (non-determinism, §3.6). + +**Judgments.** + +``` +ρ, φ ⊢ ⟨e, σ⟩ ⇓ ⟨v, ρ′, σ′⟩ e evaluates to v +ρ, φ ⊢ ⟨e, σ⟩ ⇓ err E e fails with E (per the §3.7 table) +``` + +The environment is threaded in *and out* because `Variable` writes into the +innermost scope; since binding only ever targets the innermost scope, popping +a scope restores the environment that surrounded it. Error propagation is +left-to-right: the first failing premise's error is the conclusion of the +rule (propagation rules are omitted below). + +**Rules.** + +``` +(SCALAR) ───────────────────────────────────── + ρ, φ ⊢ ⟨s, σ⟩ ⇓ ⟨coerce(s), ρ, σ⟩ coerce per §2.2 + (null: err TypeError) + + ρ, φ ⊢ ⟨e, σ⟩ ⇓ ⟨t, ρ′, σ′⟩ t ∈ Term +(URI-REF) ───────────────────────────────────── + ρ, φ ⊢ ⟨{"@id": e}, σ⟩ ⇓ ⟨uri(lex(t)), ρ′, σ′⟩ + + ρᵢ₋₁, φ ⊢ ⟨eᵢ, σᵢ₋₁⟩ ⇓ ⟨vᵢ, ρᵢ, σᵢ⟩ (i = 1…n, document order) +(OBJ) ───────────────────────────────────── + ρ₀, φ ⊢ ⟨{k₁:e₁,…,kₙ:eₙ}, σ₀⟩ ⇓ ⟨{k₁:v₁,…,kₙ:vₙ}, ρₙ, σₙ⟩ + +(DATA) as (OBJ), but only operation-call holes are evaluated (threaded + in document order); all other members are left untouched and the + value is the resulting JSON structure + + ρ₀ = ρ·∅ ρᵢ₋₁, φ ⊢ ⟨eᵢ, σᵢ₋₁⟩ ⇓ ⟨vᵢ, ρᵢ, σᵢ⟩ (i = 1…n) +(SEQ) ───────────────────────────────────── + ρ, φ ⊢ ⟨[e₁,…,eₙ], σ₀⟩ ⇓ ⟨v₁ ⧺ … ⧺ vₙ, pop(ρₙ), σₙ⟩ + + op has no quoted operands + ρᵢ₋₁, φ ⊢ ⟨eᵢ, σᵢ₋₁⟩ ⇓ ⟨vᵢ, ρᵢ, σᵢ⟩ (i = 1…n, document order) + δ_op(v₁,…,vₙ, σₙ) = (v, σ′) +(CALL) ───────────────────────────────────── + ρ₀, φ ⊢ ⟨op(k₁:e₁,…,kₙ:eₙ), σ₀⟩ ⇓ ⟨v, ρₙ, σ′⟩ + + ρ, φ ⊢ ⟨e, σ⟩ ⇓ ⟨v, ρ′, σ′⟩ +(VARIABLE) ───────────────────────────────────── + ρ, φ ⊢ ⟨Variable(name: x, value: e), σ⟩ + ⇓ ⟨unit, bind(ρ′, x, v), σ′⟩ + bind writes x ↦ v into the innermost scope (pushing one onto an + empty stack); rebinding overwrites + +(VALUE-VAR) ρ, φ ⊢ ⟨Value(name: $x), σ⟩ ⇓ ⟨lookup(ρ, x), ρ, σ⟩ + lookup searches scopes innermost→outermost; miss: err ValueError + +(VALUE-CTX) φ = (c, i, n) + ρ, φ ⊢ ⟨Value(name: x), σ⟩ ⇓ ⟨member(c, x), ρ, σ⟩ + member per §3.5, defined only for c ∈ Binding + mapping; + φ = ⊥, other item shapes, or miss: err ValueError + +(CURRENT) φ = (c, i, n) ⇒ ρ, φ ⊢ ⟨Current(), σ⟩ ⇓ ⟨c, ρ, σ⟩ +(POSITION) φ = (c, i, n) ⇒ ρ, φ ⊢ ⟨Position(), σ⟩ ⇓ ⟨int(i), ρ, σ⟩ +(LAST) φ = (c, i, n) ⇒ ρ, φ ⊢ ⟨Last(), σ⟩ ⇓ ⟨int(n), ρ, σ⟩ + each: φ = ⊥ ⇒ err ValueError; int(·) is an xsd:integer Literal + + ρ, φ ⊢ ⟨e_sel, σ⟩ ⇓ ⟨C, ρ′, σ₀⟩ items(C) = c₁ … cₙ + ρ′·∅, (cᵢ, i, n) ⊢ ⟨q, σᵢ₋₁⟩ ⇓ ⟨wᵢ, _, σᵢ⟩ (i = 1…n) +(FOREACH) ───────────────────────────────────── + ρ, φ ⊢ ⟨ForEach(select: e_sel, operation: q⟨quoted⟩), σ⟩ + ⇓ ⟨w₁ ⧺ … ⧺ wₙ, ρ′, σₙ⟩ + items(Sequence) = its elements; items(Result) = its rows in + result order; other C: err TypeError. If q is an array + [q₁,…,q_m], the premise evaluates it as a sequence within the + iteration's scope and wᵢ is its value (SEQ). + + ρ, φ ⊢ params member forms (document order) ⇓ P, ρ′, σ₀ (eager) + loop(ρ′·{P}, σ₀, 1) = ⟨w₁ ⧺ … ⧺ w_m, σ′⟩ +(ITERATE) ───────────────────────────────────── + ρ, φ ⊢ ⟨Iterate(params, operation: q⟨quoted⟩, + next-iteration: N⟨quoted⟩, break: b), σ⟩ + ⇓ ⟨w₁ ⧺ … ⧺ w_m, ρ′, σ′⟩ + + where loop(ρ_it, σ, k) is defined by: + 1. ρ_it·∅, φ ⊢ ⟨q, σ⟩ ⇓ ⟨w, ρ_b, σ₁⟩ (array operand as in FOREACH) + 2. if N is absent or k = CAP: ⟨w, σ₁⟩ + 3. ρ_b, φ ⊢ N's member forms in document order ⇓ v₁ … vⱼ, σ₂ + (each binding nᵢ ↦ vᵢ before the next evaluates) + 4. ρ_it ← ρ_it with each nᵢ ↦ vᵢ rebound in the loop scope; + the body scope is dropped + 5. if b holds of ρ_it: ⟨w, σ₂⟩ + 6. else ⟨W, σ″⟩ = loop(ρ_it, σ₂, k+1); ⟨w ⧺ W, σ″⟩ + CAP = 1000 (normative). b holds iff the lexical form of the + named loop variable equals (resp. differs from) the eagerly + evaluated comparison value; a missing variable compares as "". +``` + +**Metatheory.** Because forms are finite terms, there is no recursion, +`ForEach` iterates a *computed, finite* sequence, and `Iterate` is bounded +by its normative iteration cap, every evaluation terminates provided every +δ_op does (structural induction on forms, with (FOREACH) measured by the +size of `items(C)` and (ITERATE) by `CAP − k`). Evaluation is deterministic +up to the declared non-deterministic operators and the World's own behavior. In +(FOREACH), each iteration's environment writes are confined to its fresh +scope and results are concatenated in item order, so evaluating iterations +concurrently is observationally equivalent to evaluating them in sequence +for programs in which no iteration updates a resource that another iteration +updates or reads. §3.6 makes the first condition a rule, by making the same +target updated twice an error. The second is left to the program, as XSLT +leaves XTRE1500: a read of a resource that another iteration updates is +implementation-dependent, and the equivalence holds only for programs whose +result does not depend on it. + +## 4. Operation Catalog (normative) + +Catalog entry conventions: the *Abstract* signature is in the type language of +§1.1; *JSON args* lists the argument keys of the JSON serialization with their +expected value types after evaluation; ⟨quoted⟩ marks quoted operands (§3.3). +`Maybe τ` marks optional arguments. Effects per §3.6 are noted when not pure. + +### 4.1 Control flow, variables, context + +**ForEach** — evaluate an operation once per item of a sequence or per row of +a SPARQL result; establishes the focus *(item, position, size)* (§3.5). +``` +Abstract: (Sequence α + Result) × Operation⟨quoted⟩ → Sequence β +Python: execute_json only (interpreter-level special form) +JSON: select: Sequence α + Result · operation⟨quoted⟩: form or array of forms +``` +- Iterates a `Sequence` item-by-item, a `Result` row-by-row in result order. + Any other `select` value raises `TypeError`. +- Each iteration runs in a fresh variable scope under the focus + *(item i, i, n)*. +- If `operation` is an array, its forms evaluate in order within the + iteration's scope and the iteration's value is the array's value as a + sequence form: the concatenation of its element values (§3.2). +- The result is the concatenation of the iteration values in item order + (§3.1): a sequence-valued iteration contributes its items and a Unit-valued + one (`None`) contributes nothing, exactly as `xsl:for-each` builds its + result sequence. + +**Iterate** — stateful iteration with parameter passing between iterations, +inspired by XSLT 3.0's `xsl:iterate`; shared with the REST-VKG (XML) +serialization. +``` +Abstract: (Name ⇀ Value) × Operation⟨quoted⟩ + × Maybe (Name ⇀ Operation⟨quoted⟩) × Maybe Break → Sequence β +Python: execute_json only (interpreter-level special form) +JSON: params: Maybe object of name → form (eager) + · operation⟨quoted⟩: form or array of forms + · next-iteration⟨quoted⟩: Maybe object of name → form + · break: Maybe {name: String, equals: form} + or {name: String, not-equals: form} +``` +- `params` members are evaluated once, eagerly, in the enclosing environment + and bound as variables in a fresh loop scope (read via `$name`). +- Each iteration evaluates `operation` in a fresh scope inside the loop scope + (array operands as in `ForEach`: the concatenation of the element values). + The result is the concatenation of the iteration values, in order (§3.1). +- After the body, each `next-iteration` member is evaluated *in the + iteration's environment* — the loop parameters plus any bindings the body + made — in document order, each visible to the ones after it; the results + rebind the loop parameters for the next iteration. Without + `next-iteration`, exactly one iteration runs. +- `break` is tested after the parameters are rebound: the lexical form of + the named loop variable is compared with the eagerly evaluated `equals` + (or `not-equals`) value; a missing variable compares as the empty string. + Exactly one of `equals`/`not-equals` is required (`ValueError` otherwise). + This is the structured form of the XML serialization's + `test="$name = 'literal'"` / `!=` condition. +- The iteration count is bounded by a **normative cap of 1000**; reaching it + stops the loop (it is not an error), which keeps `Iterate` — and the + algebra — terminating. +- `Iterate` does not establish a focus; the enclosing focus, if any, remains + visible to the body. + +**Filter** — selection by position from a sequence, or by name from a row, +XPath-style. +``` +Abstract: (Sequence α + Result + Binding) × (Position + Literal) → α +Python: def execute(self, input_data: Any, expression: Any) -> Any +JSON: input: Sequence α + Result + Binding · expression: Position or Literal +``` +- With a `Position`, like a positional predicate (`$seq[2]`): 1-based; a + `Result` input is treated as its row sequence (yields a `Binding`). + Position < 1 or > length raises `ValueError`. +- With a string Literal on a `Binding`, like the lookup operator (`$row?url`): + yields the term bound to that variable name, given bare or with either of + SPARQL's variable sigils (`?url`, `$url`); a miss raises `ValueError`, as + the focus lookup of §3.5 does. +- Any other pairing — a Literal on a sequence or result, a Position on a + Binding, an expression that is neither integer nor string — raises + `TypeError`. `Filter(Filter(PUT(…), 1), "url")` is the URI a write reports + (§4.4). + +**Bindings** — project a SPARQL result to its row sequence. +``` +Abstract: Result → Sequence Binding +Python: def execute(self, table: Result) -> List[Dict[str, Node]] +JSON: table: Result +``` +- Order-preserving; an empty result yields the empty sequence. + +**Variable** — bind a name in the current scope (like `xsl:variable`). +``` +Abstract: String × Any → Unit +Python: def execute(self, name: str, value: Any, variable_stack: list) -> None +JSON: name: String (plain JSON string, not a form) · value: any form +``` +- Binds in the innermost scope (§3.4). Returns Unit — the empty sequence — and + so contributes no item wherever it stands (§3.1), as `xsl:variable` does; + the Python layer's `None` is that value. + +**Value** — read a variable (`$name`) or a focus-item member (`name`). §3.4–3.5. +``` +Abstract: String → Any +Python: def execute(self, name: str, context: Any, variable_stack: list) -> Any +JSON: name: String (plain JSON string; `$` prefix selects variable lookup) +``` +- Returns the value as-is, with no atomization or string conversion — the + semantics of `xsl:sequence`, not `xsl:value-of` (which is expressible as + `Str(Value(...))`). +- Focus-item lookup is defined for `Binding` and mapping items only (§3.5). + +**Current** — the focus item itself, per XSLT `current()`. +``` +Abstract: () → Context Python: def execute(self, current_item: Any) -> Any +JSON: (no arguments) ``` +- Raises `ValueError` when no focus is established (§3.5). -**Execute** - Execute nested operation +**Position** — the 1-based position of the focus item, per XPath +`fn:position()`. ``` -Abstract: Operation → Any -Python: def execute(self, operation: Any) -> Any +Abstract: () → Literal +Python: def execute(self, focus: Focus) -> Literal +JSON: (no arguments) ``` +- An `xsd:integer` Literal; within a focus, 1 ≤ position ≤ size. Raises + `ValueError` when no focus is established (§3.5). -**URI** - Convert term to URI reference -``` -Abstract: Term → URI -Python: def execute(self, term: rdflib.term.Node) -> rdflib.URIRef +**Last** — the size of the iterated sequence, per XPath `fn:last()`. ``` - -**ForEach** - Map operation over sequence (sequence → sequence semantics) +Abstract: () → Literal +Python: def execute(self, focus: Focus) -> Literal +JSON: (no arguments) +``` +- An `xsd:integer` Literal. Raises `ValueError` when no focus is established + (§3.5). + +### 4.2 String and term operations + +Operations in this section are named after SPARQL 1.1 / XPath F&O functions +and follow those definitions exactly — including their signatures (e.g. +`simple literal STR(literal ltrl)` / `simple literal STR(IRI rsrc)`); the +summaries below are paraphrases, and where they fall short the W3C text is +normative. A SPARQL *simple literal* is materialized as an rdflib `Literal` +with no datatype and no language tag — exactly as rdflib's own SPARQL engine +does — and under RDF 1.1 denotes the same value as the corresponding +`xsd:string` literal. + +String-compatibility rule: where an operation is documented as accepting a +*string-compatible* Literal it accepts `xsd:string` literals, language-tagged +literals, and plain literals; any other Term raises `TypeError` (use `Str` to +cast explicitly). + +**Str** — the lexical form of a Term, per SPARQL 1.1 `STR()`: +`simple literal STR(literal ltrl)` / `simple literal STR(IRI rsrc)`. +``` +Abstract: (URI + Literal) → Literal +Python: def execute(self, term: Node) -> Literal +JSON: input: URI + Literal +``` +- Returns the lexical form of a Literal, or the codepoint representation of a + URI, as a **simple literal**. As in SPARQL, the language tag is **not** + carried over. A `BNode` raises `TypeError` (a SPARQL type error), as does + any non-Term. + +**Concat** — per SPARQL 1.1 `CONCAT()`: +`string literal CONCAT(string literal ltrl1 ... string literal ltrln)`. +``` +Abstract: Sequence Literal → Literal +Python: def execute(self, inputs: List[Literal]) -> Literal +JSON: inputs: array of string-compatible Literal forms +``` +- Result kind per SPARQL: if all inputs are typed `xsd:string`, so is the + result; if all inputs carry the *same* language tag, the result carries it + too; in all other cases (including the empty input sequence) the result is + a simple literal. + +**Replace** — per SPARQL 1.1 `REPLACE()` / XPath `fn:replace`: +`string literal REPLACE(string literal arg, simple literal pattern, +simple literal replacement [, simple literal flags])`. +``` +Abstract: Literal × Literal × Literal × Maybe Literal → Literal +Python: def execute(self, input_str, pattern, replacement, flags=None) -> Literal +JSON: input: string-compatible Literal · pattern · replacement · flags: + language-tag-free string Literals (simple literals) +``` +- Per the signature, `pattern`, `replacement` and `flags` are simple + literals: a language-tagged value there raises `TypeError`. `input` may be + any string literal. +- Pattern and `flags` (`s`, `m`, `i`, `x`, `q`) per XPath `fn:replace`. In the + replacement string, `$N` references capture group *N*, `\$` is a literal + dollar, and `\\` is a literal backslash; any other use of `\` or `$` is an + error. +- Per the SPARQL string-function convention, the result is a string literal + of the same kind as `arg` (its datatype and language tag are carried over). +- Errors (`ValueError`, mirroring XPath `err:FORX000*`): invalid flags, an + invalid pattern, a pattern that matches the zero-length string, or an + invalid replacement string. + +**EncodeForURI** — percent-encode a string for use inside a URI, per SPARQL +`ENCODE_FOR_URI` / XPath `fn:encode-for-uri`. ``` -Abstract: Sequence α × Operation → Sequence β -Python: def execute(self, select_data: Union[List[Any], rdflib.query.Result], operation: Any) -> List[Any] +Abstract: Literal → Literal +Python: def execute(self, input_str: Literal) -> Literal +JSON: input: string-compatible Literal ``` +- Every character except the RFC 3986 unreserved set + (`A–Z a–z 0–9 - . _ ~`) is percent-encoded (UTF-8). Per the signature + `simple literal ENCODE_FOR_URI(string literal ltrl)`, the result is a + simple literal. -**Filter** - Filter sequences or select from results +**STRUUID** — fresh UUID string, per SPARQL `STRUUID()`. Non-deterministic. ``` -Abstract: (Sequence α × Expression → α) + (Result × Expression → Result) -Python: def execute(self, input_data: Any, expression: Any) -> Union[list, Any] +Abstract: () → Literal +Python: def execute(self) -> Literal +JSON: (no arguments) ``` +- A simple literal (per the signature `simple literal STRUUID()`) holding an + RFC 4122 version-4 UUID in lowercase hyphenated form. Successive + invocations differ. -**Bindings** - Extract binding sequence from SPARQL results +**URI** — cast a Term to a URI, like SPARQL `URI()`/`IRI()`. ``` -Abstract: Result → Sequence ResultRow -Python: def execute(self, table: rdflib.query.Result) -> List[Dict[str, Any]] +Abstract: (URI + Literal) → URI +Python: def execute(self, term: Node) -> URIRef +JSON: input: URI + Literal ``` +- A URI input is returned as-is; a Literal yields the URI of its lexical + form. A `BNode` raises `TypeError` (a blank node has no IRI). The lexical + form is *not* validated against RFC 3986; garbage in, garbage out. -### String Operations - -**Str** - Convert any term to string literal +**ResolveURI** — RFC 3986 reference resolution. ``` -Abstract: Term → Literal -Python: def execute(self, term: rdflib.term.Node) -> rdflib.Literal +Abstract: URI × Literal → URI +Python: def execute(self, base: URIRef, relative: Literal) -> URIRef +JSON: base: URI · relative: string-compatible Literal ``` +- Standard §5 resolution semantics (as by `urljoin`): if `relative` is itself + an absolute URI, the result is `relative`. -**Replace** - Replace patterns in strings using regex -``` -Abstract: Literal × Literal × Literal → Literal -Python: def execute(self, input_str: rdflib.Literal, pattern: rdflib.Literal, replacement: rdflib.Literal) -> rdflib.Literal -``` +### 4.3 SPARQL operations -**EncodeForURI** - URL-encode strings for URI usage -``` -Abstract: Literal → Literal -Python: def execute(self, input_str: Literal) -> Literal -``` +`SELECT`, `CONSTRUCT` and `DESCRIBE` run a query over a dataset given either +as `endpoint` — the URI of a SPARQL endpoint — or as `graph` — a `Graph` in +hand: the output of `GET`, `Merge` or an extraction, or an RDF data form. A +SPARQL query runs over a dataset wherever it lives, as `xsl:apply-templates` +runs over a tree in hand as readily as over `doc()`. Exactly one of the two is +given: neither raises `KeyError`, both `TypeError`. With `graph` the operation +is *pure* and local; with `endpoint` it is a *query* effect, and the response +format is negotiated transparently (SPARQL Results for `SELECT`, an RDF +serialization for `CONSTRUCT`/`DESCRIBE`) — a response that does not parse as +the negotiated format raises `ValueError` (§3.7). -**STRUUID** - Generate random UUID string +**SELECT** — execute a SPARQL SELECT query over an endpoint or a graph. ``` -Abstract: () → Literal -Python: def execute(self) -> Literal +Abstract: (URI + Graph) × Literal → Result +Python: def execute(self, source: URIRef | Graph, query: Literal) -> Result +JSON: endpoint: URI or graph: Graph (exactly one) + · query: string Literal (simple or xsd:string) ``` +- Types are validated before any network I/O. -### SPARQL Operations - -**SELECT** - Execute SPARQL SELECT query +**CONSTRUCT** — execute a SPARQL CONSTRUCT query over an endpoint or a graph. ``` -Abstract: URI × Literal → Result -Python: def execute(self, endpoint: rdflib.URIRef, query: rdflib.Literal) -> rdflib.query.Result +Abstract: (URI + Graph) × Literal → Graph +Python: def execute(self, source: URIRef | Graph, query: Literal) -> Graph +JSON: endpoint: URI or graph: Graph (exactly one) + · query: string Literal (simple or xsd:string) ``` -**CONSTRUCT** - Execute SPARQL CONSTRUCT query +**DESCRIBE** — execute a SPARQL DESCRIBE query over an endpoint or a graph. ``` -Abstract: URI × Literal → Graph -Python: def execute(self, endpoint: rdflib.URIRef, query: rdflib.Literal) -> rdflib.Graph +Abstract: (URI + Graph) × Literal → Graph +Python: def execute(self, source: URIRef | Graph, query: Literal) -> Graph +JSON: endpoint: URI or graph: Graph (exactly one) + · query: string Literal (simple or xsd:string) ``` +- What a description contains is the query processor's choice, as SPARQL 1.1 + §16.4 leaves it. -**DESCRIBE** - Execute SPARQL DESCRIBE query -``` -Abstract: URI × Literal → Graph -Python: def execute(self, endpoint: rdflib.URIRef, query: rdflib.Literal) -> rdflib.Graph -``` +In the JSON serialization, `graph` takes a `Graph` value or an RDF data form; +the form is parsed with no base IRI, so its IRIs must be absolute (§2.3). -**Substitute** - Replace variables in SPARQL queries +**Substitute** — textually substitute one SPARQL variable with a Term. ``` -Abstract: Literal × Literal × Term → Literal -Python: def execute(self, query: Literal, var: Literal, binding_value: Any) -> Literal +Abstract: Literal × Literal × (URI + Literal) → Literal +Python: def execute(self, query, var, binding_value) -> Literal +JSON: query: Literal · var: Literal (variable name, with or without `?`) + · binding: URI + Literal (Term or SPARQL JSON term form, §2.4) ``` +- Matches both `?var` and `$var` occurrences at token boundaries. A URI value + serializes as ``; a Literal as a quoted literal with its language tag + or datatype. A `BNode` value raises `TypeError` (a blank-node label in a + query is a fresh variable, not a reference — substitution would be + meaningless). +- The substitution is textual, not parse-aware; it can produce an invalid + query if `var` collides with content inside string literals of the query. + Prefer `Values` where applicable. -**Values** - Append a VALUES data block from a result set to a SPARQL query +**Values** — append a SPARQL `VALUES` data block built from a result set. ``` Abstract: Literal × Result × Maybe (Sequence Literal) → Literal -Python: def execute(self, query: Literal, data: Result, vars: Optional[List[str]] = None) -> Literal -``` - -**SPARQLString** - Generate SPARQL queries from natural language -``` -Abstract: Literal → Literal -Python: def execute(self, question: Literal) -> Literal -``` - -### HTTP Operations - -**GET** - Retrieve RDF data via HTTP GET +Python: def execute(self, query: Literal, data: Result, + vars: Optional[List[str]] = None) -> Literal +JSON: query: Literal · data: Result · vars: Maybe (array of string Literals) +``` +- Columns default to the result's variables; `vars` selects/reorders them + (names given with or without `?`). Missing values render as `UNDEF`. Terms + serialize per SPARQL syntax with correct escaping. Blank nodes raise + `ValueError` (forbidden in `VALUES`). + +**SPARQLString** — write a SPARQL query for an endpoint from a natural-language +question, via an LLM. *Non-deterministic*; external service call. It is the +`xsl:evaluate` of the algebra: the query is a string computed at run time, and +`projection` is that instruction's `as`, the shape the result must have. +``` +Abstract: URI × Literal × Maybe (Sequence Literal) × Maybe (Sequence Value) + → Literal +Python: def execute(self, endpoint: URIRef, question: Literal, + projection: Optional[List[Literal]] = None, + context: Optional[List[Any]] = None) -> Literal +JSON: endpoint: URI · question: string-compatible Literal + · projection: Maybe (array of string Literals — variable names, + given bare or with `?`/`$`) + · context: Maybe (array of forms) +``` +- The result is a simple literal holding a SPARQL 1.1 query that parses. When + `projection` is given, the query is a `SELECT` that projects every named + variable, named exactly so (it may project others as well), because the + program reads those names from its rows. +- `context` forms are evaluated eagerly, like any argument. Their values (a + `Result`'s rows, a `Graph`) are shown to the model as what the endpoint + holds, cut to a prompt-sized excerpt, before it writes the query. This is + how a program explores an endpoint it does not know (a predicate + inventory, a label lookup) with ordinary operations that are visible in + the program. An empty value (a `Result` with no rows, an empty `Graph`) is + an error (`ValueError`), raised before the model is called: what the + exploration assumed does not match the endpoint, and the query cannot be + written from it. +- An answer that does not parse, or does not have the declared projection, + goes back to the model with the reason, a bounded number of times, before + the operation fails with `ValueError`. The implementation may also put the + query's pattern to `endpoint` as an `ASK`, and send a query that matches + nothing back too. That read decides the string and is never returned; the + rows are `SELECT`'s. Because an empty answer can be the true one, a query + that still matches nothing after the last attempt is returned. +- The generated query text is not otherwise specified. + +### 4.4 Linked Data (HTTP) operations + +The Linked Data operations are RDF-specific and **symmetric**: they read and +write RDF graphs. Content negotiation is handled transparently by the +implementation — RDF media types are requested and offered; the concrete +serializations on the wire are implementation detail and never visible in +the algebra. A response that is not an RDF representation — an unsupported +media type, a missing `Content-Type`, or a body that does not parse as its +declared RDF type — raises `ValueError` (§3.7). + +`POST`, `PUT` and `PATCH` return a single-row `Result` with variables +`status` (`xsd:integer` HTTP status) and `url`. `url` is the URI of the +resource that the write produced: the response's `Location` when it has one, +as when a `POST` to a container creates a child, otherwise the effective +request URI after redirects. A relative `Location` is resolved against the +effective request URI (RFC 3986 §5). +A response outside the 2xx range is an error (`ValueError`, §3.7), as an +`xsl:result-document` that cannot be written is: the status and the server's +reason are reported, and nothing after the write runs. Transport failures +propagate per §3.7. + +Some servers apply a write as read-modify-write and require a write to an +existing resource to be conditional (`428 Precondition Required` without +one). For them, the implementation reads the resource's entity tag with +`HEAD`, sending the same `Accept` as the write since the tag names a +negotiated variant, and sends it as `If-Match`. A resource that does not +exist, or that has no tag, is written unconditionally. The tag is read +immediately before the write, not when the program last read the resource, +so this satisfies the server's precondition but does not protect against +lost updates. A write that another client makes between the `HEAD` and the +write is answered `412`, which is an error like any other non-2xx answer. +This is transport, like content negotiation, and not visible in the algebra. + +The reported `url` is reached by lookup, `Filter(Filter(PUT(…), 1), "url")` — +the first row, then its `url` (§4.1). Where the written document is wanted as +a graph, `ForEach(select: PUT(…), operation: GET(url: Value(name: url)))` +makes the row the focus and dereferences it, as `xsl:for-each` over a single +item is used to make it the context item. + +**GET** — dereference a URI to an RDF graph. *Query* effect. ``` Abstract: URI → Graph -Python: def execute(self, url: rdflib.URIRef) -> Graph +Python: def execute(self, url: URIRef) -> Graph +JSON: url: URI ``` -**POST** - Submit RDF data via HTTP POST +**POST** — append RDF data to a resource. *Update* effect. ``` Abstract: URI × Graph → Result -Python: def execute(self, url: rdflib.URIRef, data: rdflib.Graph) -> Result +Python: def execute(self, url: URIRef, data: Graph) -> Result +JSON: url: URI · data: Graph or RDF data form (parsed with base = url) ``` -**PUT** - Replace RDF data via HTTP PUT +**PUT** — replace a resource's RDF representation. *Update* effect. ``` Abstract: URI × Graph → Result -Python: def execute(self, url: rdflib.URIRef, data: rdflib.Graph) -> Result +Python: def execute(self, url: URIRef, data: Graph) -> Result +JSON: url: URI · data: Graph or RDF data form (parsed with base = url) ``` -**PATCH** - Update RDF data via HTTP PATCH with SPARQL Update +**PATCH** — apply a SPARQL Update to a resource. *Update* effect. ``` Abstract: URI × Literal → Result Python: def execute(self, url: URIRef, update: Literal) -> Result +JSON: url: URI · update: Literal (SPARQL Update string) ``` -### LinkedDataHub Operations - -**ldh-CreateContainer** - Create LinkedDataHub container document -``` -Abstract: URI × Literal × Maybe Literal × Maybe Literal → Result -Python: def execute(self, parent_uri: rdflib.URIRef, title: rdflib.Literal, slug: rdflib.Literal = None, description: rdflib.Literal = None) -> Result -``` - -**ldh-CreateItem** - Create LinkedDataHub item document -``` -Abstract: URI × Literal × Maybe Literal → Result -Python: def execute(self, container_uri: rdflib.URIRef, title: rdflib.Literal, slug: Optional[rdflib.Literal] = None) -> Result -``` - -**ldh-List** - List LinkedDataHub resources -``` -Abstract: URI × URI → List[Dict] -Python: def execute(self, url: URIRef, endpoint: URIRef) -> list[dict] -``` +### 4.5 Graph operations -**ldh-AddView** - Add view to LinkedDataHub document -``` -Abstract: URI × URI × Literal × Maybe Literal × Maybe Literal × Maybe URI → Any -Python: def execute(self, url: URIRef, query: URIRef, title: Literal, description: Literal = None, fragment: Literal = None, mode: URIRef = None) -> Any -``` - -**ldh-AddResultSetChart** - Add result set chart to LinkedDataHub document -``` -Abstract: URI × URI × Literal × URI × Literal × Literal × Maybe Literal × Maybe Literal → Any -Python: def execute(self, url: URIRef, query: URIRef, title: Literal, chart_type: URIRef, category_var_name: Literal, series_var_name: Literal, description: Literal = None, fragment: Literal = None) -> Any -``` - -**ldh-AddSelect** - Add SPARQL SELECT service to LinkedDataHub -``` -Abstract: URI × Literal × Literal × Maybe Literal × Maybe Literal × Maybe URI → Any -Python: def execute(self, url: URIRef, query: Literal, title: Literal, description: Literal = None, fragment: Literal = None, service: URIRef = None) -> Any -``` - -**ldh-AddGenericService** - Add generic SPARQL service to LinkedDataHub -``` -Abstract: URI × URI × Literal × Maybe Literal × Maybe Literal × Maybe URI × Maybe Literal × Maybe Literal → Any -Python: def execute(self, url: URIRef, endpoint: URIRef, title: Literal, description: Literal = None, fragment: Literal = None, graph_store: URIRef = None, auth_user: Literal = None, auth_pwd: Literal = None) -> Any -``` - -**ldh-AddObjectBlock** - Add object content block to LinkedDataHub document -``` -Abstract: URI × URI × Maybe Literal × Maybe Literal × Maybe Literal × Maybe URI → Any -Python: def execute(self, url: URIRef, value: URIRef, title: Literal = None, description: Literal = None, fragment: Literal = None, mode: URIRef = None) -> Any -``` - -**ldh-AddXHTMLBlock** - Add XHTML content block to LinkedDataHub document -``` -Abstract: URI × Literal × Maybe Literal × Maybe Literal × Maybe Literal → Any -Python: def execute(self, url: URIRef, value: Literal, title: Literal = None, description: Literal = None, fragment: Literal = None) -> Any -``` - -**ldh-AddFile** - Add file (binary) to LinkedDataHub document via multipart RDF/POST -``` -Abstract: URI × Literal × Literal × Maybe Literal × Maybe Literal → Any -Python: def execute(self, url: URIRef, file_path: Literal, title: Literal, description: Literal = None, content_type: Literal = None) -> Any -``` - -**ldh-RemoveBlock** - Remove content block from LinkedDataHub document -``` -Abstract: URI × Maybe URI → Any -Python: def execute(self, url: URIRef, block: URIRef = None) -> Any -``` - -**ldh-GenerateOntologyViews** - Generate LDH views (`ldh:view`) and SPIN `sp:Select` queries for each non-`owl:FunctionalProperty` `owl:ObjectProperty` in an ontology graph; `owl:DatatypeProperty` is excluded because literal values are displayed inline in LDH and do not benefit from a table view -``` -Abstract: Graph × URI × URI → Graph -Python: def execute(self, ontology: rdflib.Graph, base_uri: URIRef, service_uri: URIRef) -> rdflib.Graph -``` - -**ldh-GenerateClassContainers** - Create an LDH container per `owl:Class` in an ontology graph (each with a SPARQL service and instance-list view) -``` -Abstract: Graph × URI × URI → Result -Python: def execute(self, ontology: rdflib.Graph, parent_container: URIRef, endpoint: URIRef) -> Result -``` - -**ldh-GeneratePortal** - End-to-end portal generation; composes `ExtractOntology`, `ldh-GenerateOntologyViews`, `POST`, and `ldh-GenerateClassContainers` -``` -Abstract: URI × URI × URI → Result -Python: def execute(self, endpoint: URIRef, ontology_namespace: URIRef, parent_container: URIRef) -> Result -``` - -### Schema Operations - -**ExtractClasses** - Extract RDF classes from graph -``` -Abstract: URI → Graph -Python: def execute(self, endpoint: URIRef) -> rdflib.Graph -``` - -**ExtractDatatypeProperties** - Extract datatype properties from graph -``` -Abstract: URI → Graph -Python: def execute(self, endpoint: URIRef) -> rdflib.Graph -``` - -**ExtractObjectProperties** - Extract object properties from instance data; infers `owl:FunctionalProperty` when global max objects-per-subject = 1 (closed-world assumption over present triples, ignores formal ontology at `/ns`) -``` -Abstract: URI → Graph -Python: def execute(self, endpoint: URIRef) -> rdflib.Graph -``` - -**ExtractOntology** - Extract a full ontology (classes + datatype + object properties) from a SPARQL endpoint as a single graph -``` -Abstract: URI → Graph -Python: def execute(self, endpoint: URIRef) -> rdflib.Graph -``` - -### Utility Operations - -**Merge** - Merge multiple RDF graphs into one +**Merge** — union of graphs. ``` Abstract: Sequence Graph → Graph -Python: def execute(self, graphs: List[rdflib.Graph]) -> rdflib.Graph -``` - -**ResolveURI** - Resolve relative URI against base URI -``` -Abstract: URI × Literal → URI -Python: def execute(self, base: URIRef, relative: Literal) -> URIRef -``` - -## Type System Properties - -### Strict Type Checking -- All operations enforce strict input type checking -- TypeError raised for mismatched input types with informative messages -- No automatic type casting or conversion -- RDFLib types must match exactly as specified in signatures - -### Execution Architecture -- **execute()**: Pure functions operating on RDFLib types only -- **execute_json()**: JSON processing layer that calls execute() with type validation -- **mcp_run()**: MCP interface layer that calls execute() with plain arguments - -### Sequence Semantics -- **ForEach**: Maps from sequences to sequences (sequence → sequence) -- **Context**: In ForEach, context is the current sequence item (ResultRow for JSONResult) -- **Filter**: Can operate on both sequences and JSONResult tables -- **Automatic Application**: Single-item operations applied element-wise to sequences in execution layer - -### Variable System -- **Lexical Scoping**: Variables follow XSLT-style lexical scoping rules -- **Variable Stack**: Maintains nested scopes for variable resolution -- **Syntax**: `$variableName` for variable references, `variableName` for context access -- **Variable**: Sets variables in current scope, Variable operation manages the stack - -### Context System -- **Context Type**: `Any` - varies by operation and execution context -- **ForEach Context**: Current sequence item (ResultRow for SPARQL results) -- **Current Operation**: Returns the current context item unchanged -- **Value Operation**: Accesses both context values and variables from stack - -### URI Resolution -- **JSON-LD Parsing**: All operations that parse JSON-LD use `base` parameter to resolve relative URIs -- **Fragment URIs**: Fragment identifiers like `#service` resolve against target document URI -- **Base URI**: Set to the target document URI for correct resolution of relative references -- **Implementation**: Uses `rdflib.Graph.parse(data=json_data, format="json-ld", base=target_url)` \ No newline at end of file +Python: def execute(self, graphs: List[Graph]) -> Graph +JSON: graphs: array of Graph or RDF data forms +``` +- Set union of triples: duplicate triples collapse. Blank-node labels are + taken as-is (this is graph union, not RDF merge — graphs sharing a label + will coalesce on it). Input order is irrelevant to the result. + +### 4.6 Schema operations + +All take `endpoint`, the URI of a **SPARQL endpoint**, query the instance +data there, and return an ontology `Graph`. *Query* effect. + +``` +Abstract: URI × Maybe Result → Graph +Python: def execute(self, endpoint: URIRef, + bindings: Optional[Result] = None) -> Graph +JSON: endpoint: URI · bindings: Maybe Result +``` + +- Without `bindings`, an extraction describes the whole endpoint. With + `bindings`, it describes only the subjects in the result's `subject` + column, which are put into the extraction query as a `VALUES` block (§4.3 + `Values`). This is how a program says "these": the instances of a class, a + sample of them, or whatever an exploration found. It keeps the extraction + usable against a public endpoint, and as a `SPARQLString` context (§4.3). +- `bindings` that do not bind the variable `subject`, or that have no rows, + raise `ValueError`. A non-`Result` raises `TypeError`. +- The extractions read the default graph and every named graph, since an + endpoint (LinkedDataHub's, for one) may keep each document in a named graph + of its own. If an endpoint refuses `GRAPH` in a query, the implementation + queries the default graph only. This is transport. +- Inferences are closed-world over the triples present. + +**ExtractClasses** — classes present in the data (`owl:Class` candidates), +from `rdf:type` usage. + +**ExtractDatatypeProperties** — `owl:DatatypeProperty` candidates from +literal-valued predicates, with `rdfs:domain` when the subjects share one +class, `rdfs:range` from the literal datatypes, and a +`owl:maxQualifiedCardinality 1` restriction on the domain when no subject +has more than one value. + +**ExtractObjectProperties** — `owl:ObjectProperty` candidates from IRI- and +blank-node-valued predicates, with `rdfs:domain` when the subjects share one +class and `rdfs:range` from a sampled object's class. It infers +`owl:FunctionalProperty` when no subject has more than one object for the +property. + +**ExtractOntology** — the union of the three extractions above (classes plus +datatype and object properties) as one graph, each scoped by the same +`bindings`. + +## 5. Conformance notes + +- The JSON serialization here and the XML serialization used by REST-VKG are + two concrete syntaxes of the same abstract algebra. They share operation + names and abstract signatures, and both implement the whole catalog of §4. + The XML serialization spells the sequence form as a `Sequence` element, + since XML has no arrays, and keeps `StrUUID` as a deprecated alias of + `STRUUID`. +- Extension families (Appendix A) are named differently in the two + serializations, since JSON has no namespaces: a family's XML namespace and + prefix correspond to a fixed JSON name prefix, given in the family's + appendix. The XML executor dispatches on the expanded name, and argument + elements stay in the algebra's namespace. +- What the XML serialization still owes the spec: + - The variable environment is threaded only inside sequence constructors + (`Sequence`, and the `operation` of `ForEach` and `Iterate`), not across + the operands of a call (§3.8 CALL, OBJ). A `Variable` written as an + argument of another operation binds nothing for its sibling arguments. + - `Iterate` binds a parameter whose value is a `Result` to the lexical + form of its first row's first value, and one whose value is a one-item + sequence to that item. The algebra binds the value as it is, and + `Str(Filter(Filter(…, 1), "name"))` expresses the collapse. `Iterate` + also gives a parameter named `totalLimit` a meaning of its own (the loop + stops once its graph items hold that many distinct subjects), which no + parameter has in the algebra. + - `Concat`, `EncodeForURI` and `Replace`'s `input` accept any literal; per + SPARQL (§4.2) a non-string literal is a type error. + - `Replace` uses Java's regular-expression dialect and replacement syntax, + which accepts `${name}` where XPath raises `err:FORX0004`, rather than + XPath's. + - `ResolveURI` resolves with `java.net.URI.resolve`, which departs from + RFC 3986 §5 in known cases (an empty reference, dot segments in the + base). + - `Value` with a bare name that misses the focus item falls back to the + variables, with a deprecation warning (§3.4 keeps the two domains apart). +- An XML text argument has no integer/string distinction, so a `Filter` + expression that reads as an integer is a position. A variable whose name + is all digits, which SPARQL allows, is looked up there as `?1`. +- MCP exposure (`mcp_run`) is an interface adapter, not part of the algebra; + its plain-JSON conversions are implementation detail. + +--- + +## Appendix A — LinkedDataHub extension operations (informative) + +The `ldh-*` operations target a LinkedDataHub instance and compose the core +operations above (mostly `PUT`/`POST`/`PATCH` with LDH vocabularies). The +update operations return the single-row `Result` of §4.4 (`status`, `url`), +so they compose exactly as `PUT` does, and are subject to the same rules +(§3.6, §4.4); `ldh-List` *(query)* returns a `Result` with one row per child +document, variables `child` (URI) and `thing` (`Maybe URI`, its +`foaf:primaryTopic`). All are *update* effects unless noted. + +They are an extension family, as `ixsl:` extends XSLT: + +| Family | XML namespace | XML prefix | JSON name prefix | +|--------|---------------|------------|------------------| +| LinkedDataHub | `https://w3id.org/atomgraph/web-algebra/linkeddatahub` | `waldh` | `ldh-` | + +In XML, an operation of the family is an element in the family's namespace +whose local name is the JSON name without the prefix (`waldh:CreateItem` is +`ldh-CreateItem`). Its argument elements are in the algebra's namespace, as +every argument is: arguments are the algebra's structure, and the family +names only what is done with them. JSON has no namespaces, so the name prefix +stands for the namespace there. A prefixed name such as `waldh:CreateItem` +cannot serve instead: the JSON operation names are also the names of the +operations' MCP tools, and an MCP tool name allows only ASCII letters, digits, +`_`, `-` and `.`. The two spellings denote the same operations. + +| Operation | JSON args (`Maybe` = optional) | +|-----------|--------------------------------| +| `ldh-CreateContainer` | `parent: URI · title: Literal · slug: Maybe Literal · description: Maybe Literal` | +| `ldh-CreateItem` | `container: URI · title: Literal · slug: Maybe Literal` | +| `ldh-List` *(query)* | `url: URI · endpoint: URI` (or `base: Literal`, from which `endpoint` = `base` + `sparql`) | +| `ldh-AddFile` | `url: URI · file: Literal (path) · title: Literal · description: Maybe Literal · content_type: Maybe Literal` | +| `ldh-AddGenericService` | `url: URI · endpoint: URI · title: Literal · description/fragment: Maybe Literal · graph_store: Maybe URI · auth_user/auth_pwd: Maybe Literal` | +| `ldh-AddResultSetChart` | `url: URI · query: URI · title: Literal · chart_type: URI · category_var_name: Literal · series_var_name: Literal · description/fragment: Maybe Literal` | +| `ldh-AddSelect` | `url: URI · query: Literal · title: Literal · description/fragment: Maybe Literal · service: Maybe URI` | +| `ldh-AddConstruct` | `url: URI · query: Literal · title: Literal · description/fragment: Maybe Literal · service: Maybe URI` | +| `ldh-AddView` | `url: URI · query: URI · title: Literal · description/fragment: Maybe Literal · mode: Maybe URI` | +| `ldh-AddObjectBlock` | `url: URI · value: URI · title/description/fragment: Maybe Literal · mode: Maybe URI` | +| `ldh-AddXHTMLBlock` | `url: URI · value: Literal (XHTML) · title/description/fragment: Maybe Literal` | +| `ldh-RemoveBlock` | `url: URI · block: Maybe URI` | +| `ldh-GenerateOntologyViews` | `ontology: Graph · base_uri: URI · service_uri: URI` | +| `ldh-GenerateClassContainers` | `ontology: Graph · parent_container: URI · endpoint: URI · service_uri: Maybe URI` | +| `ldh-GeneratePortal` | `endpoint: URI · ontology_namespace: URI · parent_container: URI` | + +`ldh-AddSelect` and `ldh-AddConstruct` record a stored query (`sp:Select`, +`sp:Construct`) in the document. Two differences between the serializations +are forced by the XML one running as a service: its `AddFile` takes `file` +as a URL that the server fetches, since a local path would let a submitted +program upload any file the server can read; and it leaves out the +`Generate*` operations, which only the JSON serialization has. diff --git a/prompts/system.md b/prompts/system.md index fe2a594..ca6eeb4 100644 --- a/prompts/system.md +++ b/prompts/system.md @@ -9,6 +9,7 @@ Your output must be a **JSON-formatted structure** of operation calls, where **o - **Operations must be represented as JSON objects**. Each operation corresponds to a function call with a specific signature. - **Operations may be nested inside arguments** to indicate dependencies. - **A result can be used directly as an argument in another operation** instead of requiring explicit intermediate variables. +- **URIs are written with the URI reference form** `{"@id": "https://..."}` — an object whose only member is `@id`. A plain JSON string is **always a string literal, never a URI**. To produce a URI from a computed value, use `URI` or `ResolveURI`. - **ForEach supports executing multiple operations sequentially** when provided with a list of operations. Each operation in the list is executed for every row in the table before moving to the next row. - **Where an operation returns or expects RDF data, it is handled internally as an `rdflib.Graph`, but is represented as JSON-LD in the JSON structure.** - **SPARQL tabular data** (e.g., from `SELECT`) can be provided inline as a list of bindings, while **RDF Graph data** (e.g., from `GET`, `CONSTRUCT`, or merges) can be provided inline as JSON-LD objects. @@ -28,11 +29,17 @@ would produce this JSON output: "select": { "@op": "SELECT", "args": { - "endpoint": "https://dbpedia.org/sparql", + "endpoint": { + "@id": "https://dbpedia.org/sparql" + }, "query": { "@op": "SPARQLString", "args": { - "question": "10 biggest cities in Denmark" + "endpoint": { + "@id": "https://dbpedia.org/sparql" + }, + "question": "10 biggest cities in Denmark", + "projection": ["city", "cityName"] } } } @@ -43,7 +50,9 @@ would produce this JSON output: "url": { "@op": "ResolveURI", "args": { - "base": "http://localhost/denmark/", + "base": { + "@id": "http://localhost/denmark/" + }, "relative": { "@op": "Value", "args": { @@ -56,7 +65,7 @@ would produce this JSON output: "@op": "GET", "args": { "url": { - "@op": "Str", + "@op": "URI", "args": { "input": { "@op": "Value", @@ -233,12 +242,14 @@ Creates or replaces a document with RDF content, represented as JSON-LD. --- -## SPARQLString(question: str) -> Union[Select, Ask, Describe, Construct] +## SPARQLString(endpoint: URL, question: str, projection?: List[str], context?: List[Callable]) -> str -This function accepts a natural language question and returns a valid SPARQL query string (either `Select` or `Describe` form) that provides a result which answers the query. Uses OpenAI's API to generate a structured SPARQL query based on the provided question. +Writes a SPARQL query for the given endpoint that answers the natural language question, using OpenAI's API. Returns the query string. Use the `Select` form when you want to list resources and their property values and get a tabular result. Use the `Describe` form when you want to get RDF graph descriptions of one or more resources. -Do not return `SELECT *` or `DESCRIBE *`. The query must explicitly list all variables projected in the result. + +- `projection` names the variables the query must project. **Always give it when the result is read by name** — e.g. by `Value` inside a `ForEach` body, or by `Filter` on a row — listing exactly those names. The query is then a `SELECT` that projects them. +- `context` holds operations, run first, whose results are shown to the model as what the endpoint holds: e.g. a `SELECT` that lists the predicates of a class, or looks an entity up by its label. Use it to explore an endpoint whose vocabulary you do not know, so the query uses terms the endpoint actually has. An exploration that returns nothing is an error. ### Example JSON @@ -246,19 +257,21 @@ Do not return `SELECT *` or `DESCRIBE *`. The query must explicitly list all var { "@op": "SPARQLString", "args": { - "question": "Provide the description of the City of Copenhagen" + "endpoint": { "@id": "https://dbpedia.org/sparql" }, + "question": "The 10 biggest cities in Denmark with their names", + "projection": ["city", "cityName"] } } ``` Result: ```sparql -"DESCRIBE " +"PREFIX dbo: SELECT ?city ?cityName WHERE { ... } ORDER BY DESC(?population) LIMIT 10" ``` -## SELECT(endpoint: URL, query: Select) -> Dict +## SELECT(endpoint | graph, query: Select) -> Dict -This function queries the provided SPARQL endpoint using the provided `Select` query string. It returns a SPARQL results object with the structure `{"results": {"bindings": [...]}}` where each binding is a dictionary representing a table row. The dictionary keys correspond to variables projected by the query. +This function runs the provided `Select` query string over a SPARQL endpoint (`endpoint`, a URL) or over a graph in hand (`graph`, e.g. the result of `GET`, `Merge` or `CONSTRUCT`, or inline JSON-LD). Give exactly one of `endpoint` and `graph`. It returns a SPARQL results object with the structure `{"results": {"bindings": [...]}}` where each binding is a dictionary representing a table row. The dictionary keys correspond to variables projected by the query. Key values are also dictionaries, with `type` field indicating the type of the value (`uri`, `bnode`, or `literal`) and `value` providing the actual value. In case of language-tagged literals there is also an `xml:lang` key indicating the language code, and in case of typed literals there is a "datatype" key indicating the datatype URI. @@ -268,7 +281,7 @@ In case of language-tagged literals there is also an `xml:lang` key indicating t { "@op": "SELECT", "args": { - "endpoint": "https://dbpedia.org/sparql", + "endpoint": { "@id": "https://dbpedia.org/sparql" }, "query": "SELECT ?city ?cityName WHERE { ?city ?cityName }" } } @@ -293,9 +306,9 @@ Result (truncated for brevity): } ``` -## DESCRIBE(endpoint: URL, query: Describe) -> Graph +## DESCRIBE(endpoint | graph, query: Describe) -> Graph -This function queries the provided SPARQL endpoint using the provided `DESCRIBE` query string. It returns an RDF graph represented as JSON-LD. +This function runs the provided `DESCRIBE` query string over a SPARQL endpoint (`endpoint`) or a graph in hand (`graph`); give exactly one. It returns an RDF graph represented as JSON-LD. ### Example JSON @@ -303,7 +316,7 @@ This function queries the provided SPARQL endpoint using the provided `DESCRIBE` { "@op": "DESCRIBE", "args": { - "endpoint": "https://dbpedia.org/sparql", + "endpoint": { "@id": "https://dbpedia.org/sparql" }, "query": "DESCRIBE " } } @@ -322,9 +335,9 @@ Result (truncated) } ``` -## CONSTRUCT(endpoint: URL, query: Construct) -> Graph +## CONSTRUCT(endpoint | graph, query: Construct) -> Graph -This function queries the provided SPARQL endpoint using the provided `CONSTRUCT` query string, returning an RDF graph internally, represented as JSON-LD in the JSON structure. +This function runs the provided `CONSTRUCT` query string over a SPARQL endpoint (`endpoint`) or a graph in hand (`graph`; give exactly one), returning an RDF graph internally, represented as JSON-LD in the JSON structure. ### Example JSON @@ -332,7 +345,7 @@ This function queries the provided SPARQL endpoint using the provided `CONSTRUCT { "@op": "CONSTRUCT", "args": { - "endpoint": "https://dbpedia.org/sparql", + "endpoint": { "@id": "https://dbpedia.org/sparql" }, "query": "PREFIX dbo: CONSTRUCT { ?p ?o } WHERE { ?p ?o }" } } @@ -478,6 +491,7 @@ Executes one or more operations for each row in a SPARQL results or any sequence - If a **single operation** is provided, it is applied to each row. - If a **list of operations** is provided, they are executed sequentially for each row. +- The result is the flat sequence of every iteration's values, in row order. Iterations may run in parallel: two iterations writing the same URL is an error, so give each row its own document. --- @@ -494,7 +508,7 @@ This example performs an HTTP **GET** request for each city in the result set. "select": { "@op": "SELECT", "args": { - "endpoint": "https://dbpedia.org/sparql", + "endpoint": { "@id": "https://dbpedia.org/sparql" }, "query": "SELECT ?city WHERE { ?city a }" } }, @@ -539,7 +553,7 @@ This example performs both a **GET** request and a **POST** request for each cit "select": { "@op": "SELECT", "args": { - "endpoint": "https://dbpedia.org/sparql", + "endpoint": { "@id": "https://dbpedia.org/sparql" }, "query": "SELECT ?city WHERE { ?city a }" } }, @@ -585,6 +599,33 @@ GET("http://dbpedia.org/resource/Aarhus") POST("https://example.com/store", "http://dbpedia.org/resource/Aarhus") ``` +## Filter(input: Union[List, Result, Row], expression: Union[int, str]) -> Any + +Selects by position or by name, like XPath's `$seq[2]` and `$row?url`. + +- An integer `expression` selects the item at that 1-based position from a sequence, or the row at that position from a `SELECT` result. +- A string `expression` on a row returns the value bound to that variable name (`"url"`, `"?url"` and `"$url"` are the same). + +`PUT`, `POST` and `PATCH` return a one-row result with `status` and `url` (the created document's URL when the server says where it put it), so `Filter(Filter(, 1), "url")` is the URL a write produced. + +### Example JSON + +```json +{ + "@op": "Filter", + "args": { + "input": { + "@op": "Filter", + "args": { + "input": { "@op": "POST", "args": { "url": { "@id": "https://localhost:4443/" }, "data": { "@id": "#this", "http://purl.org/dc/terms/title": "New item" } } }, + "expression": 1 + } + }, + "expression": "url" + } +} +``` + ## Replace(input: str, pattern: str, replacement: str) -> str This function replaces occurrences of a pattern in a string with a specified replacement value. The function follows the behavior of SPARQL’s `REPLACE()`. @@ -628,28 +669,6 @@ Result: "Malm%C3%B6%20Municipality" ``` -## Execute(operation: Dict) -> Any - -This operation executes a (potentially nested) operation from its JSON representation. The operation is expected to be an instance of the Operation class. - -### Example JSON - -```json -{ - "@op": "Execute", - "args": { - "operation": { - "@op": "GET", - "args": { - "url": "http://dbpedia.org/resource/Copenhagen" - } - } - } -} -``` - -Result: Returns the result of the executed operation. - ## ldh-List(url: str, endpoint?: str, base?: str) -> List[Dict[str, Any]] Returns a list of children documents for the given LinkedDataHub URL. Requires either an endpoint or base parameter. If base is provided, the endpoint is constructed as base + "sparql". @@ -822,9 +841,74 @@ Result (example): } ``` -## ExtractClasses(endpoint: str) -> Graph +## Iterate(params: Dict, operation: Union[Callable, List[Callable]], next-iteration: Dict, break: Dict) -> List + +Stateful iteration with parameter passing between iterations, inspired by XSLT 3.0's `xsl:iterate`. Use it for cursor- or URL-driven pagination, where the next request depends on the previous response. + +- `params`: initial parameters (name → value/operation), bound as variables and read via `{"@op": "Value", "args": {"name": "$name"}}`. +- `operation`: evaluated once per iteration; may be a list (executed in order, last non-null result is the iteration's value). +- `next-iteration` (optional): name → operation, evaluated after each iteration — it sees the loop parameters *and* any `Variable` bindings the body made — and rebinds the parameters. Without it, exactly one iteration runs. +- `break` (optional): `{"name": "param", "equals": "value"}` or `{"name": "param", "not-equals": "value"}`, tested after rebinding. + +Returns the list of iteration results. The iteration count is capped at 1000. + +### Example JSON + +```json +{ + "@op": "Iterate", + "args": { + "params": { + "url": {"@id": "https://api.example.com/items"} + }, + "operation": [ + { + "@op": "Variable", + "args": {"name": "page", "value": {"@op": "GET", "args": {"url": {"@op": "URI", "args": {"input": {"@op": "Value", "args": {"name": "$url"}}}}}}} + }, + {"@op": "Value", "args": {"name": "$page"}} + ], + "next-iteration": { + "url": {"@op": "Value", "args": {"name": "$nextPageUrl"}} + }, + "break": {"name": "url", "equals": ""} + } +} +``` + +Result: a list with one RDF graph per fetched page (merge them with `Merge` if a single graph is needed). + +## Position() -> int + +Returns the 1-based position of the current iteration item, like XPath's `fn:position()`. Only meaningful inside `ForEach`, which establishes the focus (item, position, size). + +### Example JSON + +```json +{ + "@op": "Position" +} +``` + +Result (example): `2` (an `xsd:integer` literal) while processing the second row. + +## Last() -> int + +Returns the size of the sequence being iterated, like XPath's `fn:last()`. Only meaningful inside `ForEach`. Combine with `Position` for progress-style values, e.g. "item 2 of 10". + +### Example JSON + +```json +{ + "@op": "Last" +} +``` + +Result (example): `10` (an `xsd:integer` literal) while iterating ten rows. + +## ExtractClasses(endpoint: str, bindings?: Result) -> Graph -Extracts OWL classes from an RDF dataset via SPARQL endpoint. +Extracts OWL classes from an RDF dataset via SPARQL endpoint. Optional `bindings` is a `SELECT` result whose `?subject` column limits the extraction to those subjects (e.g. the instances of a class, or a sample) — use it on large public endpoints. The same applies to every Extract* operation. ### Example JSON @@ -832,14 +916,14 @@ Extracts OWL classes from an RDF dataset via SPARQL endpoint. { "@op": "ExtractClasses", "args": { - "endpoint": "https://dbpedia.org/sparql" + "endpoint": { "@id": "https://dbpedia.org/sparql" } } } ``` Result: Returns JSON-LD graph containing OWL class definitions. -## ExtractObjectProperties(endpoint: str) -> Graph +## ExtractObjectProperties(endpoint: str, bindings?: Result) -> Graph Extracts OWL object properties from an RDF dataset via SPARQL endpoint, including domain/range detection. @@ -849,14 +933,14 @@ Extracts OWL object properties from an RDF dataset via SPARQL endpoint, includin { "@op": "ExtractObjectProperties", "args": { - "endpoint": "https://dbpedia.org/sparql" + "endpoint": { "@id": "https://dbpedia.org/sparql" } } } ``` Result: Returns JSON-LD graph containing OWL object property definitions. -## ExtractDatatypeProperties(endpoint: str) -> Graph +## ExtractDatatypeProperties(endpoint: str, bindings?: Result) -> Graph Extracts OWL datatype properties from an RDF dataset via SPARQL endpoint, including datatype analysis. @@ -866,7 +950,7 @@ Extracts OWL datatype properties from an RDF dataset via SPARQL endpoint, includ { "@op": "ExtractDatatypeProperties", "args": { - "endpoint": "https://dbpedia.org/sparql" + "endpoint": { "@id": "https://dbpedia.org/sparql" } } } ``` @@ -931,7 +1015,7 @@ The block is appended as a trailing `VALUES` clause, which joins with the query' "data": { "@op": "SELECT", "args": { - "endpoint": "https://dbpedia.org/sparql", + "endpoint": { "@id": "https://dbpedia.org/sparql" }, "query": "SELECT ?city WHERE { ?city } LIMIT 2" } } @@ -1023,7 +1107,7 @@ Result: } ``` -## ExtractOntology(endpoint: str) -> Graph +## ExtractOntology(endpoint: str, bindings?: Result) -> Graph Extracts a complete ontology (classes + datatype properties + object properties) from a SPARQL endpoint as a single merged graph. Infers structure from instance data using the closed-world assumption — does not rely on a formal ontology declaration at `/ns`. Properties where the global max objects-per-subject = 1 are emitted as `owl:FunctionalProperty`. diff --git a/pyproject.toml b/pyproject.toml index 35fd510..3a5fdd3 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -1,6 +1,6 @@ [project] name = "web-algebra" -version = "1.5.0" +version = "2.0.0" description = "Composable RDF operations in JSON" readme = "README.md" license = "Apache-2.0" diff --git a/src/web_algebra/client.py b/src/web_algebra/client.py index 1a4e3d2..a1bfd02 100644 --- a/src/web_algebra/client.py +++ b/src/web_algebra/client.py @@ -1,4 +1,4 @@ -from typing import Optional, Tuple +from typing import Optional, Protocol, Tuple import hashlib import ssl import json @@ -12,6 +12,7 @@ from rdflib import Graph from rdflib.plugins.sparql.parser import parseQuery from urllib3.filepost import encode_multipart_formdata +from web_algebra.exceptions import WriteRefusedError MEDIA_TYPES = { @@ -21,6 +22,24 @@ "application/rdf+xml": "xml", } +# HTTP methods that change the resource they address. A client reports each one +# it completes to its recorder, and that report is the only source of an +# execution's `affected_documents` — derived from what was actually sent, not +# from reading the plan, so an operation that writes somewhere the plan does not +# name outright (a `ForEach` body resolving its URL per row) is still accounted +# for. +MUTATING_METHODS = frozenset({"POST", "PUT", "PATCH", "DELETE"}) + + +class WriteRecorder(Protocol): + """What a client needs of the execution it is running inside. + + Kept to one method so `client.py` stays free of any dependency on the + service layer: the CLI passes nothing and the clients record nowhere. + """ + + def record(self, method: str, url: str) -> None: ... + class HTTPRedirectHandler308(urllib.request.HTTPRedirectHandler): def redirect_request(self, req, fp, code, msg, headers, newurl): @@ -54,12 +73,63 @@ def http_error_429(self, req, fp, code, msg, hdrs): return self.parent.open(req) +def send( + opener: urllib.request.OpenerDirector, + request: urllib.request.Request, + recorder: Optional[WriteRecorder] = None, +) -> HTTPResponse: + """Open `request`; a mutating one answered outside 2xx raises + `WriteRefusedError` (formal-semantics.md §4.4), and one that succeeded is + reported to `recorder`.""" + method = request.get_method() + try: + response = opener.open(request) + except urllib.error.HTTPError as e: + if method not in MUTATING_METHODS: + raise + raise WriteRefusedError(method, request.full_url, e.code, http_reason(e)) from None + if method in MUTATING_METHODS and not 200 <= response.status < 300: + raise WriteRefusedError(method, request.full_url, response.status, response.reason) + if recorder is not None and method in MUTATING_METHODS: + recorder.record(method, written_url(response)) + return response + + +def http_reason(error: urllib.error.HTTPError) -> str: + """What a server said when it refused a request, cut to a sentence or so: + the body when it is text, otherwise the status line's reason phrase.""" + content_type = (error.headers.get("Content-Type") or "") if error.headers else "" + textual = content_type.startswith("text/") or any( + marker in content_type for marker in ("json", "xml", "n-triples", "turtle") + ) + if textual: + try: + body = " ".join(error.read().decode("utf-8", "replace").split()) + except Exception: + body = "" + if body: + return body if len(body) <= 400 else body[:400] + "…" + return str(error.reason) + + +def written_url(response: HTTPResponse) -> str: + """The URI of the resource a write produced (formal-semantics.md §4.4): + the response's `Location` when it has one — resolved against the request, + as a relative reference may be — otherwise the effective request URI.""" + location = response.headers.get("Location") if response.headers else None + if location: + return urllib.parse.urljoin(response.geturl(), location) + return response.geturl() + + class LinkedDataClient: def __init__( self, cert_pem_path: Optional[str] = None, cert_password: Optional[str] = None, verify_ssl: bool = True, + ca_bundle: Optional[str] = None, + recorder: Optional[WriteRecorder] = None, ): """ Initializes the LinkedDataClient with SSL configuration. @@ -67,9 +137,13 @@ def __init__( :param cert_pem_path: Path to the certificate .pem file (containing both private key and certificate). :param cert_password: Password for the encrypted private key in the .pem file. :param verify_ssl: Whether to verify the server's SSL certificate. Default is True. + :param ca_bundle: Path to a CA bundle that verification trusts in addition to the + system store — how a self-signed LinkedDataHub is reached with verification left on. + :param recorder: Receives every completed mutating request, or None to record nowhere. """ + self.recorder = recorder # Always create SSL context - self.ssl_context = ssl.create_default_context() + self.ssl_context = ssl.create_default_context(cafile=ca_bundle) # Load client certificate if provided if cert_pem_path and cert_password: @@ -97,6 +171,42 @@ def __init__( ) ] + def _send(self, request: urllib.request.Request) -> HTTPResponse: + """Open a request, reporting it to the recorder when it changed something. + + Reported *after* the response arrives and against the response's own URL, + so a request that raised is not recorded as a change and a redirected one + is recorded where the write actually landed. + + A write answered outside 2xx is refused (formal-semantics.md §4.4) and + raises `WriteRefusedError` with the status and the server's reason; a + read answered so is a transport failure and propagates unwrapped (§3.7). + """ + return send(self.opener, request, self.recorder) + + def conditional(self, url: str, accept: str) -> dict: + """The headers a write to `url` carries: `If-Match` with the resource's + current entity tag, when it has one (formal-semantics.md §4.4). + + A server that applies a write as read-modify-write (LinkedDataHub) + requires a write to an existing document to be conditional, and answers + 428 without. The tag is read by HEAD with the `Accept` the write itself + sends, since it names a negotiated variant. A resource that does not + exist, or cannot be read, has no tag and is written unconditionally. + """ + request = urllib.request.Request(url, headers={"Accept": accept}, method="HEAD") + try: + response = self.opener.open(request) + except Exception: + # a resource that cannot be read cannot be matched against; the + # write answers for itself + return {} + try: + etag = response.headers.get("ETag") + finally: + response.close() + return {"If-Match": etag} if etag else {} + def get(self, url: str) -> Graph: """ Fetches RDF data from the given URL and returns it as an RDFLib Graph. @@ -114,16 +224,27 @@ def get(self, url: str) -> Graph: # Read and decode the response data data = response.read().decode("utf-8") - content_type = response.headers.get("Content-Type").split(";")[0] + # Non-RDF responses are errors (formal-semantics.md §4.4): the + # Linked Data operations read and write RDF graphs only. + content_type_header = response.headers.get("Content-Type") + content_type = ( + content_type_header.split(";")[0].strip() if content_type_header else None + ) rdf_format = MEDIA_TYPES.get(content_type) if not rdf_format: raise ValueError( - f"Unsupported Content-Type: {content_type}. Supported types are: {', '.join(MEDIA_TYPES.keys())}" + f"Non-RDF response from {url}: Content-Type {content_type!r} is not " + f"an RDF media type (supported: {', '.join(MEDIA_TYPES.keys())})" ) # Parse the RDF data into an RDFLib Graph g = Graph() - g.parse(data=data, format=rdf_format, publicID=url) + try: + g.parse(data=data, format=rdf_format, publicID=url) + except Exception as e: + raise ValueError( + f"Non-RDF response from {url}: body does not parse as {content_type}: {e}" + ) from None return g def post(self, url: str, graph: Graph) -> HTTPResponse: @@ -140,11 +261,12 @@ def post(self, url: str, graph: Graph) -> HTTPResponse: "Content-Type": "application/n-triples", "Accept": "application/n-triples", } + headers.update(self.conditional(url, headers["Accept"])) request = urllib.request.Request( url, data=data.encode("utf-8"), headers=headers, method="POST" ) - return self.opener.open(request) + return self._send(request) def put(self, url: str, graph: Graph) -> HTTPResponse: """ @@ -160,11 +282,12 @@ def put(self, url: str, graph: Graph) -> HTTPResponse: "Content-Type": "application/n-triples", "Accept": "application/n-triples", } + headers.update(self.conditional(url, headers["Accept"])) request = urllib.request.Request( url, data=data.encode("utf-8"), headers=headers, method="PUT" ) - return self.opener.open(request) + return self._send(request) def delete(self, url: str) -> HTTPResponse: """ @@ -175,7 +298,7 @@ def delete(self, url: str) -> HTTPResponse: """ request = urllib.request.Request(url, method="DELETE") - return self.opener.open(request) + return self._send(request) def patch(self, url: str, sparql_update: str) -> HTTPResponse: """ @@ -189,11 +312,12 @@ def patch(self, url: str, sparql_update: str) -> HTTPResponse: "Content-Type": "application/sparql-update", "Accept": "application/n-triples", } + headers.update(self.conditional(url, headers["Accept"])) request = urllib.request.Request( url, data=sparql_update.encode("utf-8"), headers=headers, method="PATCH" ) - return self.opener.open(request) + return self._send(request) class FileClient: @@ -226,9 +350,12 @@ def __init__( cert_pem_path: Optional[str] = None, cert_password: Optional[str] = None, verify_ssl: bool = True, + ca_bundle: Optional[str] = None, + recorder: Optional[WriteRecorder] = None, ): """Initialize TLS context + opener; mirrors `LinkedDataClient.__init__`.""" - self.ssl_context = ssl.create_default_context() + self.recorder = recorder + self.ssl_context = ssl.create_default_context(cafile=ca_bundle) if cert_pem_path and cert_password: self.ssl_context.load_cert_chain( @@ -310,7 +437,7 @@ def add_file( request = urllib.request.Request( target_url, data=body, headers=headers, method="POST" ) - response = self.opener.open(request) + response = send(self.opener, request, self.recorder) return response, sha1 @@ -320,6 +447,8 @@ def __init__( cert_pem_path: Optional[str] = None, cert_password: Optional[str] = None, verify_ssl: bool = True, + ca_bundle: Optional[str] = None, + recorder: Optional[WriteRecorder] = None, ): """ Initializes the SPARQLClient with optional SSL certificate. @@ -327,9 +456,13 @@ def __init__( :param cert_pem_path: Path to .pem file containing cert+key :param cert_password: Password for the PEM file :param verify_ssl: Whether to verify server SSL certificate + :param ca_bundle: Path to an extra CA bundle verification trusts + :param recorder: Unused here — queries read; the parameter keeps the three + clients' constructor uniform so `ClientOperation` builds any of them the same way. """ + self.recorder = recorder # Always create SSL context - self.ssl_context = ssl.create_default_context() + self.ssl_context = ssl.create_default_context(cafile=ca_bundle) # Load client certificate if provided if cert_pem_path and cert_password: @@ -355,12 +488,14 @@ def __init__( ) ] - def query(self, endpoint_url: str, query_string: str) -> dict: + def query(self, endpoint_url: str, query_string: str, post: bool = False) -> dict: """ Executes a SPARQL query. Returns Graph for CONSTRUCT/DESCRIBE, Result for SELECT/ASK. :param endpoint_url: The SPARQL endpoint URL :param query_string: SPARQL query string + :param post: Send the query as a form POST (SPARQL 1.1 Protocol §2.1.2) — + for a query too long for a URL, such as one scoped by a VALUES block :return: rdflib.Graph or rdflib.query.Result """ parsed = parseQuery(query_string) @@ -373,22 +508,57 @@ def query(self, endpoint_url: str, query_string: str) -> dict: else: raise ValueError(f"Unsupported query type: {query_type}") - # Encode URL parameters - params = urllib.parse.urlencode({"query": query_string}) - url = f"{endpoint_url}?{params}" headers = {"Accept": accept} - - request = urllib.request.Request(url, headers=headers) + if post: + headers["Content-Type"] = "application/x-www-form-urlencoded" + request = urllib.request.Request( + endpoint_url, + data=urllib.parse.urlencode({"query": query_string}).encode("utf-8"), + headers=headers, + method="POST", + ) + else: + # Encode URL parameters + params = urllib.parse.urlencode({"query": query_string}) + request = urllib.request.Request(f"{endpoint_url}?{params}", headers=headers) response = self.opener.open(request) data = response.read() if accept == "application/n-triples": g = Graph() - # convert N-Triples to JSON-LD - g.parse(data=data.decode("utf-8"), format="nt") + # convert N-Triples to JSON-LD; a body that does not parse as the + # negotiated format is an error (formal-semantics.md §4.3) + try: + g.parse(data=data.decode("utf-8"), format="nt") + except Exception as e: + raise ValueError( + f"Non-RDF response from {endpoint_url}: body does not parse " + f"as N-Triples: {e}" + ) from None jsonld_str = g.serialize(format="json-ld") jsonld_data = json.loads(jsonld_str) return jsonld_data else: - # return SPARQL JSON results as a dict + # SPARQL JSON results as a dict; json.JSONDecodeError is a + # ValueError subclass, satisfying the §4.3 error contract return json.loads(data.decode("utf-8")) + + def takes_graph(self, endpoint_url: str) -> bool: + """Whether `endpoint_url` takes `GRAPH` in a query, asked once per + endpoint with `ASK { GRAPH ?g { ?s ?p ?o } }` and remembered for the + process. Blazegraph in triples mode (Wikidata's) refuses any query with + `GRAPH` as malformed: a 4xx says no, anything else says yes and leaves + the real query to report what is wrong.""" + if endpoint_url not in _TAKES_GRAPH: + try: + self.query(endpoint_url, "ASK { GRAPH ?g { ?s ?p ?o } }", post=True) + _TAKES_GRAPH[endpoint_url] = True + except urllib.error.HTTPError as e: + _TAKES_GRAPH[endpoint_url] = not 400 <= e.code < 500 + except Exception: + _TAKES_GRAPH[endpoint_url] = True + return _TAKES_GRAPH[endpoint_url] + + +# Endpoints by whether they take GRAPH in a query (`SPARQLClient.takes_graph`). +_TAKES_GRAPH: dict = {} diff --git a/src/web_algebra/client_operation.py b/src/web_algebra/client_operation.py new file mode 100644 index 0000000..ce92d08 --- /dev/null +++ b/src/web_algebra/client_operation.py @@ -0,0 +1,63 @@ +from typing import Any, ClassVar, Type + +from http.client import HTTPResponse +from rdflib import Literal, URIRef +from rdflib.namespace import XSD +from rdflib.query import Result +from web_algebra.client import LinkedDataClient, written_url +from web_algebra.focus import report_write +from web_algebra.json_result import JSONResult + + +class ClientOperation: + """Mixin that builds an operation's HTTP client from the execution's settings. + + Replaces the `model_post_init` boilerplate that the HTTP-backed operations + (GET/POST/PUT/PATCH, SELECT/CONSTRUCT/DESCRIBE, ldh-AddFile) each repeated + verbatim. Subclasses pick the client by overriding ``client_class`` + (``LinkedDataClient`` by default; ``SPARQLClient`` for the query ops, + ``FileClient`` for the multipart file op). + + All three client classes share the ``(cert_pem_path, cert_password, + verify_ssl, ca_bundle, recorder)`` constructor, so a single builder covers + them. Every value comes off ``self.settings``, which is the object one + execution owns: that is what keeps two executions running concurrently in + one process from sharing TLS material, or from reporting their writes into + each other's ``affected_documents``. + + ``verify_ssl`` defaults to off because the command line runs against + LinkedDataHub's self-signed development certificates; the HTTP service turns + it back on and names a CA bundle instead. + """ + + client_class: ClassVar[Type] = LinkedDataClient + + def model_post_init(self, __context: Any) -> None: + self.client = self.client_class( + cert_pem_path=getattr(self.settings, "cert_pem_path", None), + cert_password=getattr(self.settings, "cert_password", None), + verify_ssl=getattr(self.settings, "verify_ssl", False), + ca_bundle=getattr(self.settings, "ca_bundle", None), + recorder=getattr(self.settings, "recorder", None), + ) + + def written(self, status: int, url: str) -> Result: + """What a write answered, as its result (formal-semantics.md §4.4): the + one-row Result of `status` and `url`. The URL is reported to the + iteration gate of an enclosing ForEach, which refuses one that another + of its iterations wrote (§3.6).""" + report_write(self.context, url) + return JSONResult( + vars=["status", "url"], + bindings=[ + { + "status": Literal(status, datatype=XSD.integer), + "url": URIRef(url), + } + ], + ) + + def written_response(self, response: HTTPResponse) -> Result: + """`written` for an HTTP write response: its `Location` when it has + one, otherwise the effective request URI.""" + return self.written(response.status, written_url(response)) diff --git a/src/web_algebra/exceptions.py b/src/web_algebra/exceptions.py new file mode 100644 index 0000000..49e81f4 --- /dev/null +++ b/src/web_algebra/exceptions.py @@ -0,0 +1,66 @@ +"""Web Algebra exception hierarchy. + +`WebAlgebraError` is the base for every failure the interpreter raises about a +Web Algebra *document* — an unknown operation, an unresolved variable, a +missing iteration focus, an ill-formed value. Each subclass also inherits the +built-in exception that the specification's error table +(`formal-semantics.md` §3.7) mandates, so the normative contract — and the +existing `pytest.raises(TypeError | ValueError | KeyError)` assertions — keep +holding, while callers gain `except WebAlgebraError` to tell an ill-formed +document apart from an unrelated bug. + +Scope: these classes cover the **interpreter / composition layer** (dispatch, +evaluation, variable and focus resolution). Individual operations validate +their own argument *types* with plain `TypeError` / `ValueError` per the +spec's Strict Type Checking property; those are not reclassified here. + +Two deliberate omissions: + +- There is no wrapper for HTTP/SPARQL transport failures. Spec §3.7 pins them + to `urllib.error.HTTPError` / `URLError` propagating **unwrapped**; wrapping + would contradict the normative contract. A *write* answered outside 2xx is + not a transport failure but a refused write (§4.4), raised as + `WriteRefusedError`. +- Per-operation argument type errors stay built-in `TypeError` for the same + reason (§3.7 names them `TypeError`), and to avoid churning ~170 leaf + validation sites whose meaning is already unambiguous. +""" + + +class WebAlgebraError(Exception): + """Base for errors the Web Algebra interpreter raises about a document.""" + + +class UnknownOperationError(WebAlgebraError, ValueError): + """An `@op` names an operation that is not registered (spec §3.7).""" + + +class InvalidFormError(WebAlgebraError, TypeError): + """A JSON value is not a valid Web Algebra form — e.g. `null` (spec §2.2, §3.7).""" + + +class VariableNotFoundError(WebAlgebraError, ValueError): + """A `$name` variable lookup or a focus-item member lookup found nothing (spec §3.7).""" + + +class NoFocusError(WebAlgebraError, ValueError): + """An operation that requires an iteration focus ran outside one (spec §3.5).""" + + +class SameTargetError(WebAlgebraError, ValueError): + """Two iterations of one `ForEach` updated the same URI (spec §3.6) — the + algebra's XTDE1490.""" + + +class WriteRefusedError(ValueError): + """A write was answered outside 2xx (spec §4.4) — as an + `xsl:result-document` that cannot be written. Carries the status and the + server's reason. Not a `WebAlgebraError`: the document is well-formed, the + world refused it.""" + + def __init__(self, method: str, url: str, status: int, reason: str): + self.method = method + self.url = url + self.status = status + self.reason = reason + super().__init__(f"{method} {url} answered {status}: {reason}") diff --git a/src/web_algebra/focus.py b/src/web_algebra/focus.py new file mode 100644 index 0000000..2176f53 --- /dev/null +++ b/src/web_algebra/focus.py @@ -0,0 +1,27 @@ +from dataclasses import dataclass +from typing import Any, Callable, Optional + + +@dataclass(frozen=True) +class Focus: + """The dynamic context of a ForEach iteration (formal-semantics.md §3.5): + the current item, its 1-based position, and the iteration size — exactly + XSLT's focus triple. Established only by ForEach; accessed by Current, + Position, Last and focus-item Value lookups. + + `written` is the iteration's write gate (§3.6): every update made while + this focus is current reports the URI it wrote, and the ForEach that + established the focus refuses a URI that another of its iterations wrote. + """ + + item: Any + position: int + size: int + written: Optional[Callable[[str], None]] = None + + +def report_write(context: Any, url: str) -> None: + """Report a completed write to the iteration gate of the focus in + `context`, if there is one; outside any ForEach writes are unconstrained.""" + if isinstance(context, Focus) and context.written is not None: + context.written(url) diff --git a/src/web_algebra/operation.py b/src/web_algebra/operation.py index 2cae5b5..6ab9314 100644 --- a/src/web_algebra/operation.py +++ b/src/web_algebra/operation.py @@ -1,4 +1,5 @@ from abc import ABC, abstractmethod +from enum import Enum import json import logging from typing import Type, Dict, Optional, Any, List, ClassVar, Union @@ -8,6 +9,11 @@ from rdflib import URIRef, Literal, BNode, Graph from rdflib.namespace import XSD from rdflib.query import Result +from web_algebra.exceptions import ( + InvalidFormError, + UnknownOperationError, + VariableNotFoundError, +) # JSON-LD keyword set used to recognise a dict as RDF data (a JSON-LD @@ -17,6 +23,33 @@ _JSONLD_KEYS = ("@context", "@graph", "@id", "@type") +class OperationKind(str, Enum): + """What an operation does to the world. + + The ordering is by blast radius, so a composite form (a `ForEach` whose body + `PATCH`es) takes the kind of the worst thing it can do — `max()` over the + kinds of the operations it contains. + """ + + READ = "read" + WRITE = "write" + DESTRUCTIVE = "destructive" + + @property + def severity(self) -> int: + return _KIND_SEVERITY[self] + + def __lt__(self, other: "OperationKind") -> bool: + return self.severity < other.severity + + +_KIND_SEVERITY = { + OperationKind.READ: 0, + OperationKind.WRITE: 1, + OperationKind.DESTRUCTIVE: 2, +} + + class Operation(ABC, BaseModel): """ Abstract base class for all operations with dual execution paths: @@ -27,8 +60,18 @@ class Operation(ABC, BaseModel): """ registry: ClassVar[Dict[str, Type["Operation"]]] = {} + + #: What this operation does to the world, for the plan summary a client shows + #: before approving a document. Reading is the default because most operations + #: only compute; the ones that issue a mutating request say so themselves. + kind: ClassVar[OperationKind] = OperationKind.READ + settings: BaseSettings = Field(exclude=True) - context: Any = {} + #: The iteration focus an enclosing ForEach established, empty outside one. + #: A `default_factory`, not a literal: a bare `{}` is one dict shared by every + #: instance ever constructed, so two executions in one process — which is what + #: the HTTP service runs — would see each other's focus. + context: Any = Field(default_factory=dict) model_config = ConfigDict(extra="allow") @@ -53,7 +96,7 @@ def execute(self, *args) -> Union[Node, Result, Graph]: @abstractmethod def execute_json( - self, arguments: dict, variable_stack: list = [] + self, arguments: dict, variable_stack: list = None ) -> Union[Node, Result, Graph]: """JSON execution: processes JSON args, returns RDFLib objects""" pass @@ -80,10 +123,14 @@ def process_json( cls, settings: BaseSettings, json_data: Any, - context: dict = {}, - variable_stack: list = [], + context: dict = None, + variable_stack: list = None, ) -> Any: """Class method for processing JSON with @op structures""" + if context is None: + context = {} + if variable_stack is None: + variable_stack = [] if isinstance(json_data, dict): if "@op" in json_data: op_name = json_data["@op"] @@ -91,7 +138,7 @@ def process_json( operation_cls = cls.get(op_name) if not operation_cls: - raise ValueError(f"Unknown operation: {op_name}") + raise UnknownOperationError(f"Unknown operation: {op_name}") operation = operation_cls(settings=settings, context=context) result = operation.execute_json(op_args, variable_stack) @@ -99,6 +146,18 @@ def process_json( # Return RDFLib objects as-is for operation chaining return result + # URI reference form (formal-semantics.md §2.2 rule 2): an object + # whose ONLY member is `@id` evaluates to a URI. This is JSON-LD's + # node-reference syntax — a bare node reference carries no triples, + # so reusing it as the URI form is unambiguous. Inside an RDF data + # form this rule never fires: `_resolve_jsonld` walks those without + # re-entering this dispatch for non-`@op` objects. + if set(json_data.keys()) == {"@id"}: + inner = cls.process_json( + settings, json_data["@id"], context, variable_stack + ) + return URIRef(str(inner)) + # JSON-LD shape recognition — a dict carrying any JSON-LD reserved # key is RDF data (a JSON-LD document or fragment), not generic # JSON to recurse into. It may still embed `@op` operations or @@ -128,25 +187,47 @@ def process_json( } elif isinstance(json_data, list): - # For sequential operations, share variable stack to allow accumulation - results = [] - current_stack = variable_stack.copy() - for item in json_data: - result = cls.process_json(settings, item, context, current_stack) - results.append(result) - return results + # Sequence form (formal-semantics.md §3.2): elements evaluate in + # order in a fresh variable scope, so a Variable bound in step N is + # visible to steps N+1.. and to nested forms, and goes out of scope + # when the sequence ends. + variable_stack.append({}) + try: + items: list = [] + for item in json_data: + cls.concatenate( + items, + cls.process_json(settings, item, context, variable_stack), + ) + return items + finally: + variable_stack.pop() else: # Convert plain values to RDFLib terms return cls.json_to_rdflib(json_data) + @staticmethod + def concatenate(items: list, value: Any) -> None: + """Append `value` to `items` as XDM concatenation does + (formal-semantics.md §3.1): a sequence contributes its items, Unit + (`None`) contributes none, anything else — a `Result` included — is + one item.""" + if value is None: + return + if isinstance(value, list): + for item in value: + Operation.concatenate(items, item) + return + items.append(value) + @classmethod def _resolve_jsonld( cls, settings: BaseSettings, json_data: Any, - context: dict = {}, - variable_stack: list = [], + context: dict = None, + variable_stack: list = None, ) -> Any: """Resolve embedded `@op` nodes inside a JSON-LD document in place. @@ -174,19 +255,6 @@ def _resolve_jsonld( else: return json_data - @staticmethod - def _serialize_for_json_context(obj) -> Any: - """Convert RDFLib objects to appropriate format for JSON consumption""" - if isinstance(obj, (URIRef, Literal, BNode)): - return str(obj) # Convert RDFLib terms to strings for JSON-LD - elif hasattr(obj, "to_json") and callable(obj.to_json): - return obj.to_json() # Convert Result to SPARQL JSON format - elif isinstance(obj, Graph): - # Keep graphs as-is for now - they'll be serialized by HTTP operations - return obj - else: - return obj - # Variable stack management methods def push_variable_scope(self, variable_stack: list): """Create a new variable scope (like entering a new XSLT template).""" @@ -208,7 +276,7 @@ def get_variable(self, name: str, variable_stack: list) -> Any: for scope in reversed(variable_stack): if name in scope: return scope[name] - raise ValueError(f"Variable '{name}' not found") + raise VariableNotFoundError(f"Variable '{name}' not found") # Conversion helpers between different formats @staticmethod @@ -244,6 +312,9 @@ def to_graph(data: Any, *, base: Optional[str] = None) -> Graph: @staticmethod def json_to_rdflib(data) -> Node: """Convert JSON/binding objects to RDFLib terms""" + if data is None: + # formal-semantics.md §2.2: null is not a valid form. + raise InvalidFormError("null is not a valid Web Algebra form") if isinstance(data, dict) and "type" in data and "value" in data: # SPARQL binding object - values may have been processed to RDFLib terms type_str = str(data["type"]) # Convert potential Literal to string @@ -275,12 +346,13 @@ def json_to_rdflib(data) -> Node: elif isinstance(data, str): # Plain string → always convert to string literal return Literal(data, datatype=XSD.string) + elif isinstance(data, bool): + # bool before int — bool is an int subclass in Python + return Literal(data, datatype=XSD.boolean) elif isinstance(data, int): return Literal(data, datatype=XSD.integer) elif isinstance(data, float): return Literal(data, datatype=XSD.double) - elif isinstance(data, bool): - return Literal(data, datatype=XSD.boolean) else: # Default: convert to string literal return Literal(str(data), datatype=XSD.string) @@ -291,15 +363,26 @@ def plain_to_rdflib(value: Any) -> Node: if isinstance(value, str): # Plain string → always convert to string literal return Literal(value, datatype=XSD.string) + elif isinstance(value, bool): + # bool before int — bool is an int subclass in Python + return Literal(value, datatype=XSD.boolean) elif isinstance(value, int): return Literal(value, datatype=XSD.integer) elif isinstance(value, float): return Literal(value, datatype=XSD.double) - elif isinstance(value, bool): - return Literal(value, datatype=XSD.boolean) else: return Literal(str(value), datatype=XSD.string) + @staticmethod + def is_string_literal(term: Any) -> bool: + """True for a language-tag-free string literal — a SPARQL simple + literal or its RDF 1.1 equivalent, an xsd:string literal.""" + return ( + isinstance(term, Literal) + and term.language is None + and (term.datatype is None or term.datatype == XSD.string) + ) + @staticmethod def to_string_literal(term: Node) -> Literal: """Convert Literal terms to string-compatible literals, following SPARQL semantics""" diff --git a/src/web_algebra/operations/bindings.py b/src/web_algebra/operations/bindings.py index ca221fe..85686a1 100644 --- a/src/web_algebra/operations/bindings.py +++ b/src/web_algebra/operations/bindings.py @@ -1,5 +1,4 @@ -from typing import Any, List, Dict -from mcp import types +from typing import List, Dict from web_algebra.operation import Operation from rdflib.query import Result from rdflib.term import Node @@ -35,7 +34,7 @@ def execute(self, table: Result) -> List[Dict[str, Node]]: return table.bindings def execute_json( - self, arguments: dict, variable_stack: list = [] + self, arguments: dict, variable_stack: list = None ) -> List[Dict[str, Node]]: """JSON execution: process arguments with strict type checking""" # Process table @@ -48,12 +47,3 @@ def execute_json( ) return self.execute(table_data) - - def mcp_run(self, arguments: dict, context: Any = None) -> Any: - """MCP execution: plain args → plain results""" - # For MCP, just return summary - return [ - types.TextContent( - type="text", text="Extracted bindings from SPARQL results" - ) - ] diff --git a/src/web_algebra/operations/current.py b/src/web_algebra/operations/current.py index 6f890a0..ea2fca7 100644 --- a/src/web_algebra/operations/current.py +++ b/src/web_algebra/operations/current.py @@ -1,5 +1,6 @@ from typing import Any -from mcp import types +from web_algebra.exceptions import NoFocusError +from web_algebra.focus import Focus from web_algebra.operation import Operation @@ -27,14 +28,20 @@ def execute(self, current_item: Any) -> Any: """Pure function: return current sequence item""" return current_item - def execute_json(self, arguments: dict, variable_stack: list = []) -> Any: - """JSON execution: return current context item""" - if self.context is None: - raise ValueError("Current operation requires context") + def execute_json(self, arguments: dict, variable_stack: list = None) -> Any: + """JSON execution: return the current focus item""" + # The focus is established by ForEach (formal-semantics.md §3.5); + # Current yields its item. + if isinstance(self.context, Focus): + return self.execute(self.context.item) + + # No focus established — the interpreter's default context is an + # empty dict. + if self.context is None or ( + isinstance(self.context, dict) and not self.context + ): + raise NoFocusError( + "Current requires an iteration focus (only ForEach establishes one)" + ) return self.execute(self.context) - - def mcp_run(self, arguments: dict, context: Any = None) -> Any: - """MCP execution: plain args → plain results""" - # For MCP, we just return a placeholder since context handling is JSON-specific - return [types.TextContent(type="text", text="Current context accessed")] diff --git a/src/web_algebra/operations/execute.py b/src/web_algebra/operations/execute.py deleted file mode 100644 index 69c962c..0000000 --- a/src/web_algebra/operations/execute.py +++ /dev/null @@ -1,56 +0,0 @@ -from typing import Any -from mcp import types -from web_algebra.operation import Operation - - -class Execute(Operation): - """ - Execute a (potentially nested) operation from its JSON representation. - """ - - @classmethod - def description(cls) -> str: - return "Executes a (potentially nested) operation from its JSON representation. The operation is expected to be an instance of the Operation class." - - @classmethod - def inputSchema(cls) -> dict: - """ - Returns the JSON schema of the operation's input arguments. - """ - return { - "type": "object", - "properties": { - "operation": { - "type": "object", - "description": "An instance of an Operation to execute.", - "properties": { - "type": { - "type": "string", - "description": "Type of the operation", - }, - }, - "required": ["type"], - } - }, - "required": ["operation"], - } - - def execute(self, operation: Any) -> Any: - """Pure function: execute operation with RDFLib terms""" - if not isinstance(operation, dict) or "@op" not in operation: - raise TypeError( - f"Execute.execute expects operation dict with '@op' key, got {type(operation)}" - ) - - # Delegate to Operation.process_json for nested operation execution - return Operation.process_json(self.settings, operation, self.context) - - def execute_json(self, arguments: dict, variable_stack: list = []) -> Any: - """JSON execution: pass raw operation to execute""" - # Don't process the operation argument - execute() expects raw operation dict - return self.execute(arguments["operation"]) - - def mcp_run(self, arguments: dict, context: Any = None) -> Any: - """MCP execution: plain args → plain results""" - result = self.execute(arguments["operation"]) - return [types.TextContent(type="text", text=str(result))] diff --git a/src/web_algebra/operations/filter.py b/src/web_algebra/operations/filter.py index 8c66ce4..27c95e0 100644 --- a/src/web_algebra/operations/filter.py +++ b/src/web_algebra/operations/filter.py @@ -1,119 +1,130 @@ -from typing import Any, Union -from mcp import types +from typing import Any, Mapping +from rdflib import Literal +from rdflib.term import Node +from rdflib.namespace import XSD +from rdflib.query import Result, ResultRow +from web_algebra.exceptions import VariableNotFoundError from web_algebra.operation import Operation class Filter(Operation): """ - Filters SPARQL results using filter expressions, similar to XSLT predicates. - Currently supports positional filtering (e.g., [1], [2]) with future extensibility. + Selection by position from a sequence, or by name from a row, XPath-style + (formal-semantics.md §4.1): `$seq[2]` and `$row?url`. """ @classmethod def description(cls) -> str: - return """Filters SPARQL results using filter expressions, similar to XSLT predicates. - - Currently supports positional access: - - Numeric values select by position (1-based, following XSLT convention) - - Returns the filtered SPARQL results with matching bindings - - Future versions may support field-based and complex filter expressions.""" + return """Selects an item by position, or a value by name, XPath-style. + + - An integer expression selects by 1-based position from a sequence or + from a SPARQL result's rows (a result yields a row): like `$seq[2]`. + - A string expression on a row (a SPARQL binding) yields the term bound + to that variable name, given bare or as ?name / $name: like `$row?url`. + + Filter(Filter(PUT(...), 1), "url") is the URI a write reports.""" @classmethod def inputSchema(cls) -> dict: return { "type": "object", "properties": { - "input": {"description": "SPARQL results to filter."}, + "input": { + "description": "A sequence, a SPARQL result, or one row of a result." + }, "expression": { - "description": "Filter expression. Supports positional integers (1-based index) and will support more expression types in the future." + "description": "An integer position (1-based) on a sequence or result, or a variable name on a row." }, }, "required": ["input", "expression"], "additionalProperties": False, } - def execute(self, input_data: Any, expression: Any) -> Union[list, Any]: - """Pure function: filter any iterable with filter expression""" - # Convert any iterable to list for processing - if hasattr(input_data, "__iter__"): - items = list(input_data) - else: - raise TypeError(f"Filter expects iterable input, got {type(input_data)}") - - # Handle different expression types - if isinstance(expression, int): - # Positional filtering (current implementation) - filtered_items = self._apply_positional_filter(items, expression) - else: - # Future: could support other expression types (boolean expressions, etc.) - raise NotImplementedError( - f"Filter expression type {type(expression)} not yet supported" + def execute(self, input_data: Any, expression: Any) -> Any: + """Pure function: select by position or look up by name""" + position = self._position(expression) + name = self._name(expression) + if position is None and name is None: + raise TypeError( + f"Filter expects an integer position or a string name, got {expression!r}" ) - # Return single item directly if only one result (XSLT semantics) - if len(filtered_items) == 1: - return filtered_items[0] - return filtered_items - - def execute_json( - self, arguments: dict, variable_stack: list = [] - ) -> Union[list, Any]: - """JSON execution: process arguments with support for both Result and sequence""" - # Process input - input_data = Operation.process_json( - self.settings, arguments["input"], self.context, variable_stack - ) - - # Process expression - keep as plain value if it's already an int - expression_arg = arguments["expression"] - if isinstance(expression_arg, int): - expression_data = expression_arg - else: - # If it's an operation, process it - expression_data = Operation.process_json( - self.settings, expression_arg, self.context, variable_stack - ) - if not isinstance(expression_data, int): + if self._is_binding(input_data): + if name is None: raise TypeError( - f"Filter operation expects 'expression' to be int, got {type(expression_data)}" + "Filter: a position selects from a sequence; on a row, name the variable to look up" ) + return self._lookup(input_data, name) - return self.execute(input_data, expression_data) - - def _apply_positional_filter(self, bindings: list, position: int) -> list: - """ - Apply positional filter (1-based indexing like XSLT). + if isinstance(input_data, Result): + items = list(input_data) + elif isinstance(input_data, list): + items = input_data + else: + raise TypeError( + f"Filter expects a sequence, a Result or a Binding as input, got {type(input_data).__name__}" + ) - :param bindings: List of RDFLib term bindings - :param position: 1-based position to select - :return: List containing single binding at the specified position - """ + if position is None: + raise TypeError( + "Filter: a name looks a variable up on a row; on a sequence or result, give a position" + ) if position < 1: raise ValueError("Position must be >= 1 (XSLT-style 1-based indexing)") - - if position > len(bindings): + if position > len(items): raise ValueError( - f"Position {position} exceeds number of bindings ({len(bindings)})" + f"Position {position} exceeds number of items ({len(items)})" ) + return items[position - 1] - # Convert to 0-based index for Python list access and return as list - return [bindings[position - 1]] - - def mcp_run(self, arguments: dict, context: Any = None) -> Any: - """MCP execution: plain args → plain results""" - # Convert plain args to RDFLib terms - input_json = arguments["input"] - from web_algebra.json_result import JSONResult - - input_result = JSONResult.from_json(input_json) - expression = arguments["expression"] - - result = self.execute(input_result, expression) - - # Return summary for MCP - return [ - types.TextContent( - type="text", text=f"Filtered to {len(result.bindings)} result(s)" + def execute_json(self, arguments: dict, variable_stack: list = None) -> Any: + """JSON execution: evaluate both operands, then select""" + input_data = Operation.process_json( + self.settings, arguments["input"], self.context, variable_stack + ) + expression = Operation.process_json( + self.settings, arguments["expression"], self.context, variable_stack + ) + return self.execute(input_data, expression) + + @staticmethod + def _position(expression: Any): + """The expression as a position: an xsd:integer Literal (what a JSON + integer coerces to, §2.2) or a plain int from the pure layer.""" + if isinstance(expression, bool): + return None + if isinstance(expression, int): + return expression + if isinstance(expression, Literal) and expression.datatype == XSD.integer: + return int(expression) + return None + + @staticmethod + def _name(expression: Any): + """The expression as a variable name, without SPARQL's `?`/`$` sigil.""" + if isinstance(expression, Literal): + if not Operation.is_string_literal(expression): + return None + expression = str(expression) + # URIRef and BNode are str subclasses, but not names + elif not isinstance(expression, str) or isinstance(expression, Node): + return None + return expression[1:] if expression[:1] in ("?", "$") else expression + + @staticmethod + def _is_binding(value: Any) -> bool: + return isinstance(value, (ResultRow, Mapping)) + + @staticmethod + def _lookup(row: Any, name: str) -> Any: + if isinstance(row, ResultRow): + value = row.asdict().get(name) + names = list(row.labels) + else: + value = row.get(name) + names = list(row.keys()) + if value is None: + raise VariableNotFoundError( + f"Filter: the row has no binding for {name}; it binds {names}" ) - ] + return value diff --git a/src/web_algebra/operations/for_each.py b/src/web_algebra/operations/for_each.py index 4f47ce0..d0d1921 100644 --- a/src/web_algebra/operations/for_each.py +++ b/src/web_algebra/operations/for_each.py @@ -1,6 +1,7 @@ -from typing import Any, List, Union +from typing import Any, Callable, Dict, List, Union import logging -from mcp import types +from web_algebra.exceptions import SameTargetError +from web_algebra.focus import Focus, report_write from web_algebra.operation import Operation from rdflib.query import Result @@ -18,7 +19,7 @@ def description(cls) -> str: - Sequences: Each item becomes the context for the operation - Result (SPARQL results): Each result row (ResultRow) becomes the context - Returns a sequence of operation results.""" + Returns the concatenation of the iteration values, in item order.""" @classmethod def inputSchema(cls) -> dict: @@ -36,15 +37,21 @@ def inputSchema(cls) -> dict: def execute( self, select_data: Union[List[Any], Result], operation: Any ) -> List[Any]: - """Pure function: apply operation to each item in sequence or SPARQL results""" - # This is complex because we need to execute operations with context - # For now, this will be handled in execute_json + """Interpreter-level special form — no pure form (formal-semantics.md §4.1). + + `ForEach` evaluates a *quoted* operand once per item under a per-item + focus and variable scope; that requires the interpreter, so it has no + pure `execute()` and lives entirely in `execute_json`. + """ raise NotImplementedError( - "ForEach pure function needs operation execution context" + "ForEach is an interpreter-level special form (formal-semantics.md " + "§4.1); use execute_json" ) - def execute_json(self, arguments: dict, variable_stack: list = []) -> List[Any]: + def execute_json(self, arguments: dict, variable_stack: list = None) -> List[Any]: """JSON execution: apply operations to each item in sequence or SPARQL results""" + if variable_stack is None: + variable_stack = [] # Get the select data (sequence or Result) select_data = Operation.process_json( self.settings, arguments["select"], self.context, variable_stack @@ -76,39 +83,61 @@ def execute_json(self, arguments: dict, variable_stack: list = []) -> List[Any]: f"ForEach expects 'select' to be sequence (list) or Result, got {type(select_data)}" ) - results = [] - for item in items: + results: List[Any] = [] + size = len(items) + # §3.6: the iterations' writes go through one gate, which refuses a + # URI that two iterations write (XSLT's XTDE1490). Within one + # iteration the sequence form orders the writes, so repeats are fine. + writers: Dict[str, int] = {} + for position, item in enumerate(items, start=1): logging.info("Processing item: %s", item) - # Handle list of operations or single operation - if isinstance(operation, list): - # Execute operations in sequence, with item as context - last_result = None + # The focus (item, position, size) per formal-semantics.md §3.5, + # accessed by Current/Position/Last and focus-item Value lookups. + focus = Focus( + item=item, + position=position, + size=size, + written=self._gate(writers, position), + ) - for op in operation: - result = Operation.process_json( - self.settings, op, context=item, variable_stack=variable_stack + # Each iteration runs in a fresh variable scope + # (formal-semantics.md §3.4): bindings made inside one iteration + # do not leak into the next. + variable_stack.append({}) + try: + # An array operand is a sequence constructor evaluated within + # the iteration's scope; either way the iteration's value is + # concatenated into the result (§3.1, §4.1). + forms = operation if isinstance(operation, list) else [operation] + for form in forms: + Operation.concatenate( + results, + Operation.process_json( + self.settings, + form, + context=focus, + variable_stack=variable_stack, + ), ) - if result is not None: - last_result = result - - # Only collect the last non-None result - if last_result is not None: - results.append(last_result) - else: - # Single operation - result = Operation.process_json( - self.settings, - operation, - context=item, - variable_stack=variable_stack, - ) - # Only collect non-None results - if result is not None: - results.append(result) + finally: + variable_stack.pop() return results - def mcp_run(self, arguments: dict, context: Any = None) -> Any: - """MCP execution: plain args → plain results""" - return [types.TextContent(type="text", text="ForEach operation completed")] + def _gate(self, writers: Dict[str, int], position: int) -> Callable[[str], None]: + """The write gate of iteration `position`. An update inside a nested + ForEach counts for every enclosing iteration too, so the report is + passed on to the gate of the focus this ForEach runs under.""" + + def written(url: str) -> None: + other = writers.setdefault(url, position) + if other != position: + raise SameTargetError( + f"ForEach wrote {url} from iterations {other} and {position}: " + "two iterations updating one URI are an error " + "(formal-semantics.md §3.6)" + ) + report_write(self.context, url) + + return written diff --git a/src/web_algebra/operations/iterate.py b/src/web_algebra/operations/iterate.py new file mode 100644 index 0000000..e501fac --- /dev/null +++ b/src/web_algebra/operations/iterate.py @@ -0,0 +1,184 @@ +from typing import Any, Callable, ClassVar, List, Optional +import logging + +from web_algebra.operation import Operation + + +class Iterate(Operation): + """ + Stateful iteration with parameter passing between iterations, inspired by + XSLT 3.0's xsl:iterate. Shared with the REST-VKG (Java/XML) Web Algebra. + """ + + # Normative iteration cap (formal-semantics.md §4.1): reaching it stops + # the loop — it is not an error — and keeps the algebra terminating. + MAX_ITERATIONS: ClassVar[int] = 1000 + + @classmethod + def description(cls) -> str: + return """Stateful iteration with parameter passing between iterations, inspired by XSLT 3.0's xsl:iterate. + + `params` are evaluated once and bound as variables (read via $name). + Each iteration evaluates `operation`; afterwards the `next-iteration` + members are evaluated in the iteration's environment and rebind the + parameters. `break` compares a loop variable against a value and + stops the loop when it matches. Without `next-iteration`, exactly one + iteration runs. Returns the sequence of iteration values. Useful for + cursor- or URL-driven pagination.""" + + @classmethod + def inputSchema(cls) -> dict: + return { + "type": "object", + "properties": { + "params": { + "type": "object", + "description": "Initial parameters: name → form, evaluated once in the enclosing environment and bound as loop variables", + }, + "operation": { + "description": "Operation(s) evaluated each iteration (quoted)" + }, + "next-iteration": { + "type": "object", + "description": "name → form, evaluated after each iteration in the iteration's environment to rebind the loop parameters (quoted)", + }, + "break": { + "type": "object", + "description": "Loop exit test: {name, equals} or {name, not-equals}; the named loop variable's lexical form is compared with the value", + }, + }, + "required": ["operation"], + } + + def execute(self, *args) -> Any: + raise NotImplementedError( + "Iterate is an interpreter-level special form (formal-semantics.md " + "§4.1); use execute_json" + ) + + def execute_json(self, arguments: dict, variable_stack: list = None) -> List[Any]: + """JSON execution: run the parameterized loop (formal-semantics.md + §3.8 (ITERATE)).""" + if variable_stack is None: + variable_stack = [] + + params_arg = arguments.get("params", {}) + if not isinstance(params_arg, dict): + raise TypeError( + f"Iterate expects 'params' to be an object of name → form, got {type(params_arg)}" + ) + operation = arguments["operation"] # quoted + next_arg = arguments.get("next-iteration") # quoted + if next_arg is not None and not isinstance(next_arg, dict): + raise TypeError( + f"Iterate expects 'next-iteration' to be an object of name → form, got {type(next_arg)}" + ) + break_test = self._parse_break(arguments.get("break"), variable_stack) + + # params: eager, evaluated in the enclosing environment + params = { + name: Operation.process_json( + self.settings, form, self.context, variable_stack + ) + for name, form in params_arg.items() + } + + results: List[Any] = [] + # The loop scope holds the iteration parameters across iterations. + variable_stack.append(dict(params)) + try: + iteration = 0 + while iteration < self.MAX_ITERATIONS: + iteration += 1 + logging.info("Iterate: iteration %d", iteration) + + # Fresh body scope per iteration (§3.4); the enclosing focus, + # if any, remains visible (Iterate establishes none). + variable_stack.append({}) + try: + # The iteration's value is concatenated into the result + # (§3.1): a sequence contributes its items, Unit none. + self._evaluate_body(operation, variable_stack, results) + + if next_arg is None: + # No next-iteration: exactly one iteration. + break + + # next-iteration members: evaluated in the iteration's + # environment (loop params + body bindings), in document + # order, each visible to the ones after it. + new_params = {} + for name, form in next_arg.items(): + new_value = Operation.process_json( + self.settings, form, self.context, variable_stack + ) + new_params[name] = new_value + self.set_variable(name, new_value, variable_stack) + finally: + variable_stack.pop() + + # Rebind the loop parameters for the next iteration. + variable_stack[-1].update(new_params) + + if break_test is not None and break_test(): + logging.info("Iterate: break condition met after %d iterations", iteration) + break + else: + logging.warning( + "Iterate: reached the iteration cap (%d)", self.MAX_ITERATIONS + ) + finally: + variable_stack.pop() + + return results + + def _evaluate_body( + self, operation: Any, variable_stack: list, results: List[Any] + ) -> None: + """Evaluate the quoted operation (form or array of forms, the latter + a sequence constructor as in ForEach) and concatenate its value into + `results`.""" + forms = operation if isinstance(operation, list) else [operation] + for form in forms: + Operation.concatenate( + results, + Operation.process_json( + self.settings, form, self.context, variable_stack + ), + ) + + def _parse_break( + self, break_arg: Any, variable_stack: list + ) -> Optional[Callable[[], bool]]: + """Build the break test: lexical-form (in)equality between a loop + variable and an eagerly evaluated comparison value. A missing + variable compares as the empty string (REST-VKG parity).""" + if break_arg is None: + return None + if not isinstance(break_arg, dict) or "name" not in break_arg: + raise TypeError( + "Iterate 'break' must be an object with 'name' and exactly one of 'equals'/'not-equals'" + ) + has_equals = "equals" in break_arg + has_not_equals = "not-equals" in break_arg + if has_equals == has_not_equals: + raise ValueError( + "Iterate 'break' requires exactly one of 'equals'/'not-equals'" + ) + name = break_arg["name"] + comparison_form = break_arg["equals" if has_equals else "not-equals"] + comparison = Operation.process_json( + self.settings, comparison_form, self.context, variable_stack + ) + comparison_lex = str(comparison) + + def test() -> bool: + try: + value = self.get_variable(name, variable_stack) + except ValueError: + value = "" + lexical = str(value) if value is not None else "" + matches = lexical == comparison_lex + return matches if has_equals else not matches + + return test diff --git a/src/web_algebra/operations/last.py b/src/web_algebra/operations/last.py new file mode 100644 index 0000000..2d0f03a --- /dev/null +++ b/src/web_algebra/operations/last.py @@ -0,0 +1,35 @@ +from rdflib import Literal +from rdflib.namespace import XSD + +from web_algebra.exceptions import NoFocusError +from web_algebra.focus import Focus +from web_algebra.operation import Operation + + +class Last(Operation): + """ + Returns the size of the iterated sequence, per XPath's fn:last(). + """ + + @classmethod + def description(cls) -> str: + return """Returns the size of the sequence being iterated, per XPath's fn:last(). + + Only meaningful inside ForEach, which establishes the focus + (item, position, size).""" + + @classmethod + def inputSchema(cls) -> dict: + return {"type": "object", "properties": {}, "additionalProperties": False} + + def execute(self, focus: Focus) -> Literal: + """Pure function: focus → iteration size as xsd:integer""" + if not isinstance(focus, Focus): + raise NoFocusError( + "Last requires an iteration focus (only ForEach establishes one)" + ) + return Literal(focus.size, datatype=XSD.integer) + + def execute_json(self, arguments: dict, variable_stack: list = None) -> Literal: + """JSON execution: read the size from the current focus""" + return self.execute(self.context) diff --git a/src/web_algebra/operations/linked_data/get.py b/src/web_algebra/operations/linked_data/get.py index 3b3d16d..154afdd 100644 --- a/src/web_algebra/operations/linked_data/get.py +++ b/src/web_algebra/operations/linked_data/get.py @@ -3,23 +3,17 @@ from typing import Any from mcp import types from web_algebra.mcp_tool import MCPTool +from web_algebra.client_operation import ClientOperation from web_algebra.operation import Operation -from web_algebra.client import LinkedDataClient -class GET(Operation, MCPTool): +class GET(ClientOperation, Operation, MCPTool): """ Retrieves RDF data from a named graph using HTTP GET. The URL serves as both the resource identifier and the named graph address in systems with direct graph identification. Returns the RDF graph (describing the resource at that URL) as JSON-LD. """ - def model_post_init(self, __context: Any) -> None: - self.client = LinkedDataClient( - cert_pem_path=getattr(self.settings, "cert_pem_path", None), - cert_password=getattr(self.settings, "cert_password", None), - verify_ssl=False, # Optionally disable SSL verification - ) @classmethod def description(cls) -> str: @@ -53,7 +47,7 @@ def execute(self, url: URIRef) -> Graph: return graph - def execute_json(self, arguments: dict, variable_stack: list = []) -> Graph: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Graph: """JSON execution: process arguments and delegate to execute()""" url_data = Operation.process_json( self.settings, arguments["url"], self.context, variable_stack diff --git a/src/web_algebra/operations/linked_data/patch.py b/src/web_algebra/operations/linked_data/patch.py index f969a11..1d2a423 100644 --- a/src/web_algebra/operations/linked_data/patch.py +++ b/src/web_algebra/operations/linked_data/patch.py @@ -1,15 +1,15 @@ -from typing import Any +from typing import ClassVar, Any import logging from rdflib import URIRef, Literal from rdflib.namespace import XSD from mcp import types from web_algebra.mcp_tool import MCPTool -from web_algebra.operation import Operation +from web_algebra.client_operation import ClientOperation +from web_algebra.operation import Operation, OperationKind from rdflib.query import Result -from web_algebra.client import LinkedDataClient -class PATCH(Operation, MCPTool): +class PATCH(ClientOperation, Operation, MCPTool): """ Updates RDF data in a named graph using HTTP PATCH with SPARQL Update. The URL serves as both the resource identifier and the named graph address in systems with direct graph identification. @@ -20,12 +20,9 @@ class PATCH(Operation, MCPTool): Note: This operation does not return the updated graph, it only confirms the success of the operation. """ - def model_post_init(self, __context: Any) -> None: - self.client = LinkedDataClient( - cert_pem_path=getattr(self.settings, "cert_pem_path", None), - cert_password=getattr(self.settings, "cert_password", None), - verify_ssl=False, # Optionally disable SSL verification - ) + # a SPARQL Update whose DELETE selects triples by pattern - what it removes is not visible in the plan, only the target is + kind: ClassVar[OperationKind] = OperationKind.DESTRUCTIVE + @classmethod def description(cls) -> str: @@ -76,20 +73,9 @@ def execute(self, url: URIRef, update: Literal) -> Result: response = self.client.patch(url_str, update_str) logging.info("PATCH operation status: %s", response.status) - # Return SPARQL results format - from web_algebra.json_result import JSONResult - - return JSONResult( - vars=["status", "url"], - bindings=[ - { - "status": Literal(response.status, datatype=XSD.integer), - "url": URIRef(response.geturl()), - } - ], - ) + return self.written_response(response) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Result: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Result: """JSON execution: process arguments and call pure function""" # Process URL url_data = Operation.process_json( diff --git a/src/web_algebra/operations/linked_data/post.py b/src/web_algebra/operations/linked_data/post.py index 40242c5..c77b8d4 100644 --- a/src/web_algebra/operations/linked_data/post.py +++ b/src/web_algebra/operations/linked_data/post.py @@ -1,15 +1,14 @@ -from typing import Any +from typing import ClassVar, Any import logging -from rdflib import URIRef, Graph, Literal -from rdflib.namespace import XSD +from rdflib import URIRef, Graph from mcp import types from web_algebra.mcp_tool import MCPTool -from web_algebra.operation import Operation +from web_algebra.client_operation import ClientOperation +from web_algebra.operation import Operation, OperationKind from rdflib.query import Result -from web_algebra.client import LinkedDataClient -class POST(Operation, MCPTool): +class POST(ClientOperation, Operation, MCPTool): """ Creates or appends RDF data to a named graph using HTTP POST. The URL serves as both the resource identifier and the named graph address in systems with direct graph identification. @@ -18,12 +17,9 @@ class POST(Operation, MCPTool): Note: This operation does not return the updated graph, it only confirms the success of the operation. """ - def model_post_init(self, __context: Any) -> None: - self.client = LinkedDataClient( - cert_pem_path=getattr(self.settings, "cert_pem_path", None), - cert_password=getattr(self.settings, "cert_password", None), - verify_ssl=False, # Optionally disable SSL verification - ) + # appends the supplied triples to a document; nothing already there is removed + kind: ClassVar[OperationKind] = OperationKind.WRITE + @classmethod def description(cls) -> str: @@ -66,20 +62,9 @@ def execute(self, url: URIRef, data: Graph) -> Result: response = self.client.post(url_str, data) logging.info("POST operation status: %s", response.status) - # Return SPARQL results format - from web_algebra.json_result import JSONResult - - return JSONResult( - vars=["status", "url"], - bindings=[ - { - "status": Literal(response.status, datatype=XSD.integer), - "url": URIRef(response.geturl()), - } - ], - ) + return self.written_response(response) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Result: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Result: """JSON execution: process arguments with strict type checking""" # Process URL url_data = Operation.process_json( diff --git a/src/web_algebra/operations/linked_data/put.py b/src/web_algebra/operations/linked_data/put.py index 16a5760..17f0075 100644 --- a/src/web_algebra/operations/linked_data/put.py +++ b/src/web_algebra/operations/linked_data/put.py @@ -1,15 +1,14 @@ -from typing import Any +from typing import ClassVar, Any import logging -from rdflib import URIRef, Graph, Literal -from rdflib.namespace import XSD +from rdflib import URIRef, Graph from mcp import types from web_algebra.mcp_tool import MCPTool -from web_algebra.operation import Operation +from web_algebra.client_operation import ClientOperation +from web_algebra.operation import Operation, OperationKind from rdflib.query import Result -from web_algebra.client import LinkedDataClient -class PUT(Operation, MCPTool): +class PUT(ClientOperation, Operation, MCPTool): """ Replaces RDF data in a named graph using HTTP PUT. The URL serves as both the resource identifier and the named graph address in systems with direct graph identification. @@ -20,12 +19,9 @@ class PUT(Operation, MCPTool): Note: This operation does not return the updated graph, it only confirms the success of the operation. """ - def model_post_init(self, __context: Any) -> None: - self.client = LinkedDataClient( - cert_pem_path=getattr(self.settings, "cert_pem_path", None), - cert_password=getattr(self.settings, "cert_password", None), - verify_ssl=False, # Optionally disable SSL verification - ) + # writes the supplied graph at a URL the plan names outright, so the blast radius is visible in the plan itself + kind: ClassVar[OperationKind] = OperationKind.WRITE + @classmethod def description(cls) -> str: @@ -66,20 +62,9 @@ def execute(self, url: URIRef, data: Graph) -> Result: response = self.client.put(url_str, data) logging.info("PUT operation status: %s", response.status) - # Return SPARQL results format - from web_algebra.json_result import JSONResult - - return JSONResult( - vars=["status", "url"], - bindings=[ - { - "status": Literal(response.status, datatype=XSD.integer), - "url": URIRef(response.geturl()), - } - ], - ) + return self.written_response(response) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Result: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Result: """JSON execution: process arguments with strict type checking""" # Process URL url_data = Operation.process_json( diff --git a/src/web_algebra/operations/linkeddatahub/add_construct.py b/src/web_algebra/operations/linkeddatahub/add_construct.py new file mode 100644 index 0000000..79e0c2c --- /dev/null +++ b/src/web_algebra/operations/linkeddatahub/add_construct.py @@ -0,0 +1,228 @@ +from typing import Any, Optional +import logging +from rdflib import Literal, URIRef +from rdflib.namespace import XSD +from web_algebra.operation import Operation +from web_algebra.operations.linked_data.post import POST + + +class AddConstruct(POST): + @classmethod + def name(cls): + return "ldh-AddConstruct" + + @classmethod + def description(cls) -> str: + return """Creates a SPARQL CONSTRUCT query resource in a document. + + This tool creates a SPARQL CONSTRUCT query that can be executed and referenced within LinkedDataHub. + The query can be used for charts, views, or other data processing operations. + + This tool: + - Creates a sp:Construct resource with the SPARQL query text + - Posts the new CONSTRUCT query resource to the target document + - Supports optional title, description, fragment identifier, and service URI + Note: the service URI is _not_ the SPARQL endpoint URL but an instance of `sd:Service` that describes the SPARQL service capabilities (and contains the endpoint URL). + The service URI can be used to reference the SPARQL service in other operations. + """ + + @classmethod + def inputSchema(cls) -> dict: + return { + "type": "object", + "properties": { + "url": { + "type": "string", + "description": "The URI of the document to append the CONSTRUCT query to.", + }, + "query": { + "type": "string", + "description": "The SPARQL CONSTRUCT query string.", + }, + "title": { + "type": "string", + "description": "Title of the CONSTRUCT query.", + }, + "description": { + "type": "string", + "description": "Optional description of the CONSTRUCT query.", + }, + "fragment": { + "type": "string", + "description": "Optional fragment identifier for the query URI (e.g., 'my-query' creates #my-query).", + }, + "service": { + "type": "string", + "description": "Optional URI of the SPARQL service/endpoint specific to this query. Note: the service URI is _not_ the SPARQL endpoint URL but an instance of `sd:Service` that describes the SPARQL service capabilities (and contains the endpoint URL).", + }, + }, + "required": ["url", "query", "title"], + } + + def execute_json(self, arguments: dict[str, str], variable_stack: list = None) -> Any: + """JSON execution: process arguments and delegate to execute()""" + # Process required arguments + url_data = Operation.process_json( + self.settings, arguments["url"], self.context, variable_stack + ) + if not isinstance(url_data, URIRef): + raise TypeError( + f"AddConstruct operation expects 'url' to be URIRef, got {type(url_data)}" + ) + + query_data = Operation.process_json( + self.settings, arguments["query"], self.context, variable_stack + ) + query_literal = self.to_string_literal(query_data) + + title_data = Operation.process_json( + self.settings, arguments["title"], self.context, variable_stack + ) + title_literal = self.to_string_literal(title_data) + + # Process optional arguments + description_literal = None + if "description" in arguments: + description_data = Operation.process_json( + self.settings, arguments["description"], self.context, variable_stack + ) + description_literal = self.to_string_literal(description_data) + + fragment_literal = None + if "fragment" in arguments: + fragment_data = Operation.process_json( + self.settings, arguments["fragment"], self.context, variable_stack + ) + fragment_literal = self.to_string_literal(fragment_data) + + service_uri = None + if "service" in arguments: + service_data = Operation.process_json( + self.settings, arguments["service"], self.context, variable_stack + ) + if not isinstance(service_data, URIRef): + raise TypeError( + f"AddConstruct operation expects 'service' to be URIRef, got {type(service_data)}" + ) + service_uri = service_data + + return self.execute( + url_data, + query_literal, + title_literal, + description_literal, + fragment_literal, + service_uri, + ) + + def execute( + self, + url: URIRef, + query: Literal, + title: Literal, + description: Optional[Literal] = None, + fragment: Optional[Literal] = None, + service: Optional[URIRef] = None, + ) -> Any: + """Pure function: create SPARQL CONSTRUCT query with RDFLib terms""" + if not isinstance(url, URIRef): + raise TypeError( + f"AddConstruct.execute expects url to be URIRef, got {type(url)}" + ) + if not Operation.is_string_literal(query): + raise TypeError( + f"AddConstruct.execute expects query to be string Literal, got {type(query)}" + ) + if not Operation.is_string_literal(title): + raise TypeError( + f"AddConstruct.execute expects title to be string Literal, got {type(title)}" + ) + if description is not None and ( + not Operation.is_string_literal(description) + ): + raise TypeError( + f"AddConstruct.execute expects description to be string Literal, got {type(description)}" + ) + if fragment is not None and ( + not Operation.is_string_literal(fragment) + ): + raise TypeError( + f"AddConstruct.execute expects fragment to be string Literal, got {type(fragment)}" + ) + if service is not None and not isinstance(service, URIRef): + raise TypeError( + f"AddConstruct.execute expects service to be URIRef, got {type(service)}" + ) + + url_str = str(url) + query_str = str(query) + title_str = str(title) + description_str = str(description) if description else None + fragment_str = str(fragment) if fragment else None + service_str = str(service) if service else None + + logging.info( + "Creating CONSTRUCT query for document <%s> with title '%s'", + url_str, + title_str, + ) + + # Create subject URI (fragment or blank node) - matching shell script logic + if fragment_str: + subject_id = f"#{fragment_str}" # relative URI that will be resolved against the request URI + else: + subject_id = "_:subject" + + # Build JSON-LD structure for the CONSTRUCT query - matching shell script output + data = { + "@context": { + "ldh": "https://w3id.org/atomgraph/linkeddatahub#", + "dct": "http://purl.org/dc/terms/", + "sp": "http://spinrdf.org/sp#", + }, + "@id": subject_id, + "@type": "sp:Construct", + "dct:title": title_str, + "sp:text": query_str, + } + + # Add optional properties - matching shell script conditional logic + if service_str: + data["ldh:service"] = {"@id": service_str} + + if description_str: + data["dct:description"] = description_str + + logging.info(f"Posting CONSTRUCT query with JSON-LD data: {data}") + + # Convert the JSON-LD content to a Graph and POST to the target URI + graph = self.to_graph(data, base=url_str) + return super().execute(url, graph) + + def mcp_run(self, arguments: dict, context: Any = None) -> Any: + """MCP execution: plain args → plain results""" + from mcp import types + + # Convert plain arguments to RDFLib terms + url = URIRef(arguments["url"]) + query = Literal(arguments["query"], datatype=XSD.string) + title = Literal(arguments["title"], datatype=XSD.string) + + description = None + if "description" in arguments: + description = Literal(arguments["description"], datatype=XSD.string) + + fragment = None + if "fragment" in arguments: + fragment = Literal(arguments["fragment"], datatype=XSD.string) + + service = None + if "service" in arguments: + service = URIRef(arguments["service"]) + + # Call pure function + result = self.execute(url, query, title, description, fragment, service) + + # Return status for MCP response + status_binding = result.bindings[0]["status"] + return [types.TextContent(type="text", text=f"CONSTRUCT query added - status: {status_binding}")] diff --git a/src/web_algebra/operations/linkeddatahub/add_file.py b/src/web_algebra/operations/linkeddatahub/add_file.py index b10bd47..bd5e058 100644 --- a/src/web_algebra/operations/linkeddatahub/add_file.py +++ b/src/web_algebra/operations/linkeddatahub/add_file.py @@ -1,4 +1,4 @@ -from typing import Any, Optional +from typing import Any, Optional, ClassVar, Type import logging import mimetypes import urllib.parse @@ -10,12 +10,12 @@ from rdflib.query import Result from web_algebra.client import FileClient -from web_algebra.json_result import JSONResult from web_algebra.mcp_tool import MCPTool -from web_algebra.operation import Operation +from web_algebra.client_operation import ClientOperation +from web_algebra.operation import Operation, OperationKind -class AddFile(Operation, MCPTool): +class AddFile(ClientOperation, Operation, MCPTool): """RDF/POST a file to a LinkedDataHub document, returning the minted upload URI. The file's RDF description (`nfo:FileDataObject` + filename + MIME type + @@ -29,12 +29,11 @@ class AddFile(Operation, MCPTool): `FileClient` instance instead of inheriting `LinkedDataClient` plumbing. """ - def model_post_init(self, __context: Any) -> None: - self.client = FileClient( - cert_pem_path=getattr(self.settings, "cert_pem_path", None), - cert_password=getattr(self.settings, "cert_password", None), - verify_ssl=False, - ) + # uploads a file and appends its description to the target document + kind: ClassVar[OperationKind] = OperationKind.WRITE + + client_class: ClassVar[Type] = FileClient + @classmethod def name(cls): @@ -155,17 +154,9 @@ def execute( logging.info("AddFile status %s → <%s>", response.status, file_uri) - return JSONResult( - vars=["status", "url"], - bindings=[ - { - "status": Literal(response.status, datatype=XSD.integer), - "url": URIRef(file_uri), - } - ], - ) + return self.written(response.status, file_uri) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Result: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Result: """JSON execution: process arguments with strict type checking.""" url_data = Operation.process_json( self.settings, arguments["url"], self.context, variable_stack diff --git a/src/web_algebra/operations/linkeddatahub/add_generic_service.py b/src/web_algebra/operations/linkeddatahub/add_generic_service.py index 403943d..96da8f6 100644 --- a/src/web_algebra/operations/linkeddatahub/add_generic_service.py +++ b/src/web_algebra/operations/linkeddatahub/add_generic_service.py @@ -62,7 +62,7 @@ def inputSchema(cls) -> dict: "required": ["url", "endpoint", "title"], } - def execute_json(self, arguments: dict, variable_stack: list = []) -> Any: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Any: """JSON execution: process arguments and delegate to execute()""" # Process required arguments url_data = Operation.process_json( @@ -157,18 +157,18 @@ def execute( raise TypeError( f"AddGenericService.execute expects endpoint to be URIRef, got {type(endpoint)}" ) - if not isinstance(title, Literal) or title.datatype != XSD.string: + if not Operation.is_string_literal(title): raise TypeError( f"AddGenericService.execute expects title to be string Literal, got {type(title)}" ) if description is not None and ( - not isinstance(description, Literal) or description.datatype != XSD.string + not Operation.is_string_literal(description) ): raise TypeError( f"AddGenericService.execute expects description to be string Literal, got {type(description)}" ) if fragment is not None and ( - not isinstance(fragment, Literal) or fragment.datatype != XSD.string + not Operation.is_string_literal(fragment) ): raise TypeError( f"AddGenericService.execute expects fragment to be string Literal, got {type(fragment)}" @@ -178,13 +178,13 @@ def execute( f"AddGenericService.execute expects graph_store to be URIRef, got {type(graph_store)}" ) if auth_user is not None and ( - not isinstance(auth_user, Literal) or auth_user.datatype != XSD.string + not Operation.is_string_literal(auth_user) ): raise TypeError( f"AddGenericService.execute expects auth_user to be string Literal, got {type(auth_user)}" ) if auth_pwd is not None and ( - not isinstance(auth_pwd, Literal) or auth_pwd.datatype != XSD.string + not Operation.is_string_literal(auth_pwd) ): raise TypeError( f"AddGenericService.execute expects auth_pwd to be string Literal, got {type(auth_pwd)}" diff --git a/src/web_algebra/operations/linkeddatahub/add_result_set_chart.py b/src/web_algebra/operations/linkeddatahub/add_result_set_chart.py index eac7a3a..e86e737 100644 --- a/src/web_algebra/operations/linkeddatahub/add_result_set_chart.py +++ b/src/web_algebra/operations/linkeddatahub/add_result_set_chart.py @@ -86,7 +86,7 @@ def inputSchema(cls) -> dict: ], } - def execute_json(self, arguments: dict[str, str], variable_stack: list = []) -> Any: + def execute_json(self, arguments: dict[str, str], variable_stack: list = None) -> Any: """JSON execution: process arguments and delegate to execute()""" # Process required arguments url_data = Operation.process_json( @@ -174,7 +174,7 @@ def execute( raise TypeError( f"AddResultSetChart.execute expects query to be URIRef, got {type(query)}" ) - if not isinstance(title, Literal) or title.datatype != XSD.string: + if not Operation.is_string_literal(title): raise TypeError( f"AddResultSetChart.execute expects title to be string Literal, got {type(title)}" ) @@ -183,27 +183,25 @@ def execute( f"AddResultSetChart.execute expects chart_type to be URIRef, got {type(chart_type)}" ) if ( - not isinstance(category_var_name, Literal) - or category_var_name.datatype != XSD.string + not Operation.is_string_literal(category_var_name) ): raise TypeError( f"AddResultSetChart.execute expects category_var_name to be string Literal, got {type(category_var_name)}" ) if ( - not isinstance(series_var_name, Literal) - or series_var_name.datatype != XSD.string + not Operation.is_string_literal(series_var_name) ): raise TypeError( f"AddResultSetChart.execute expects series_var_name to be string Literal, got {type(series_var_name)}" ) if description is not None and ( - not isinstance(description, Literal) or description.datatype != XSD.string + not Operation.is_string_literal(description) ): raise TypeError( f"AddResultSetChart.execute expects description to be string Literal, got {type(description)}" ) if fragment is not None and ( - not isinstance(fragment, Literal) or fragment.datatype != XSD.string + not Operation.is_string_literal(fragment) ): raise TypeError( f"AddResultSetChart.execute expects fragment to be string Literal, got {type(fragment)}" diff --git a/src/web_algebra/operations/linkeddatahub/add_select.py b/src/web_algebra/operations/linkeddatahub/add_select.py index d1bf1c4..512fa8b 100644 --- a/src/web_algebra/operations/linkeddatahub/add_select.py +++ b/src/web_algebra/operations/linkeddatahub/add_select.py @@ -59,7 +59,7 @@ def inputSchema(cls) -> dict: "required": ["url", "query", "title"], } - def execute_json(self, arguments: dict[str, str], variable_stack: list = []) -> Any: + def execute_json(self, arguments: dict[str, str], variable_stack: list = None) -> Any: """JSON execution: process arguments and delegate to execute()""" # Process required arguments url_data = Operation.process_json( @@ -129,22 +129,22 @@ def execute( raise TypeError( f"AddSelect.execute expects url to be URIRef, got {type(url)}" ) - if not isinstance(query, Literal) or query.datatype != XSD.string: + if not Operation.is_string_literal(query): raise TypeError( f"AddSelect.execute expects query to be string Literal, got {type(query)}" ) - if not isinstance(title, Literal) or title.datatype != XSD.string: + if not Operation.is_string_literal(title): raise TypeError( f"AddSelect.execute expects title to be string Literal, got {type(title)}" ) if description is not None and ( - not isinstance(description, Literal) or description.datatype != XSD.string + not Operation.is_string_literal(description) ): raise TypeError( f"AddSelect.execute expects description to be string Literal, got {type(description)}" ) if fragment is not None and ( - not isinstance(fragment, Literal) or fragment.datatype != XSD.string + not Operation.is_string_literal(fragment) ): raise TypeError( f"AddSelect.execute expects fragment to be string Literal, got {type(fragment)}" diff --git a/src/web_algebra/operations/linkeddatahub/add_view.py b/src/web_algebra/operations/linkeddatahub/add_view.py index 44eab5a..b0ddaf5 100644 --- a/src/web_algebra/operations/linkeddatahub/add_view.py +++ b/src/web_algebra/operations/linkeddatahub/add_view.py @@ -76,7 +76,7 @@ def inputSchema(cls) -> dict: "required": ["url", "query", "title"], } - def execute_json(self, arguments: dict[str, str], variable_stack: list = []) -> Any: + def execute_json(self, arguments: dict[str, str], variable_stack: list = None) -> Any: """JSON execution: process arguments and delegate to execute()""" # Process required arguments url_data = Operation.process_json( @@ -153,18 +153,18 @@ def execute( raise TypeError( f"AddView.execute expects query to be URIRef, got {type(query)}" ) - if not isinstance(title, Literal) or title.datatype != XSD.string: + if not Operation.is_string_literal(title): raise TypeError( f"AddView.execute expects title to be string Literal, got {type(title)}" ) if description is not None and ( - not isinstance(description, Literal) or description.datatype != XSD.string + not Operation.is_string_literal(description) ): raise TypeError( f"AddView.execute expects description to be string Literal, got {type(description)}" ) if fragment is not None and ( - not isinstance(fragment, Literal) or fragment.datatype != XSD.string + not Operation.is_string_literal(fragment) ): raise TypeError( f"AddView.execute expects fragment to be string Literal, got {type(fragment)}" diff --git a/src/web_algebra/operations/linkeddatahub/content/add_object_block.py b/src/web_algebra/operations/linkeddatahub/content/add_object_block.py index e281efd..62d7e31 100644 --- a/src/web_algebra/operations/linkeddatahub/content/add_object_block.py +++ b/src/web_algebra/operations/linkeddatahub/content/add_object_block.py @@ -77,7 +77,7 @@ def inputSchema(cls) -> dict: "required": ["url", "value"], } - def execute_json(self, arguments: dict[str, str], variable_stack: list = []) -> Any: + def execute_json(self, arguments: dict[str, str], variable_stack: list = None) -> Any: """JSON execution: process arguments and delegate to execute()""" # Process required arguments url_data = Operation.process_json( @@ -157,19 +157,19 @@ def execute( f"AddObjectBlock.execute expects value to be URIRef, got {type(value)}" ) if title is not None and ( - not isinstance(title, Literal) or title.datatype != XSD.string + not Operation.is_string_literal(title) ): raise TypeError( f"AddObjectBlock.execute expects title to be string Literal, got {type(title)}" ) if description is not None and ( - not isinstance(description, Literal) or description.datatype != XSD.string + not Operation.is_string_literal(description) ): raise TypeError( f"AddObjectBlock.execute expects description to be string Literal, got {type(description)}" ) if fragment is not None and ( - not isinstance(fragment, Literal) or fragment.datatype != XSD.string + not Operation.is_string_literal(fragment) ): raise TypeError( f"AddObjectBlock.execute expects fragment to be string Literal, got {type(fragment)}" diff --git a/src/web_algebra/operations/linkeddatahub/content/add_xhtml_block.py b/src/web_algebra/operations/linkeddatahub/content/add_xhtml_block.py index 9f96992..014d86a 100644 --- a/src/web_algebra/operations/linkeddatahub/content/add_xhtml_block.py +++ b/src/web_algebra/operations/linkeddatahub/content/add_xhtml_block.py @@ -68,7 +68,7 @@ def inputSchema(cls) -> dict: "required": ["url", "value"], } - def execute_json(self, arguments: dict[str, str], variable_stack: list = []) -> Any: + def execute_json(self, arguments: dict[str, str], variable_stack: list = None) -> Any: """JSON execution: process arguments and delegate to execute()""" # Process required arguments url_data = Operation.process_json( @@ -140,19 +140,19 @@ def execute( f"AddXHTMLBlock.execute expects value to be XMLLiteral, got {type(value)}" ) if title is not None and ( - not isinstance(title, Literal) or title.datatype != XSD.string + not Operation.is_string_literal(title) ): raise TypeError( f"AddXHTMLBlock.execute expects title to be string Literal, got {type(title)}" ) if description is not None and ( - not isinstance(description, Literal) or description.datatype != XSD.string + not Operation.is_string_literal(description) ): raise TypeError( f"AddXHTMLBlock.execute expects description to be string Literal, got {type(description)}" ) if fragment is not None and ( - not isinstance(fragment, Literal) or fragment.datatype != XSD.string + not Operation.is_string_literal(fragment) ): raise TypeError( f"AddXHTMLBlock.execute expects fragment to be string Literal, got {type(fragment)}" diff --git a/src/web_algebra/operations/linkeddatahub/content/generate_class_containers.py b/src/web_algebra/operations/linkeddatahub/content/generate_class_containers.py index 0051aa0..3f5185b 100644 --- a/src/web_algebra/operations/linkeddatahub/content/generate_class_containers.py +++ b/src/web_algebra/operations/linkeddatahub/content/generate_class_containers.py @@ -1,8 +1,9 @@ +from typing import ClassVar import logging from rdflib import URIRef, Literal, Namespace, Graph from rdflib.namespace import RDF, RDFS, XSD, DCTERMS from rdflib.query import Result -from web_algebra.operation import Operation +from web_algebra.operation import Operation, OperationKind from web_algebra.operations.linkeddatahub.create_item import CreateItem from web_algebra.operations.linked_data.post import POST from web_algebra.operations.linkeddatahub.content.add_object_block import AddObjectBlock @@ -21,6 +22,9 @@ class GenerateClassContainers(Operation): This operation orchestrates actual HTTP operations to set up the portal structure. """ + # composes writes internally rather than through nested forms, so it declares its own kind + kind: ClassVar[OperationKind] = OperationKind.WRITE + @classmethod def name(cls): return "ldh-GenerateClassContainers" @@ -216,7 +220,7 @@ def _build_view_graph(self, view_uri: URIRef, class_local: str, query_uri: URIRe return g - def execute_json(self, arguments: dict, variable_stack: list = []) -> Result: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Result: """JSON execution: process arguments with type checking""" # Process ontology graph — accept a Graph (e.g. from CONSTRUCT) or a # JSON-LD document, converting the latter to a Graph diff --git a/src/web_algebra/operations/linkeddatahub/content/generate_ontology_views.py b/src/web_algebra/operations/linkeddatahub/content/generate_ontology_views.py index f202c16..3815dad 100644 --- a/src/web_algebra/operations/linkeddatahub/content/generate_ontology_views.py +++ b/src/web_algebra/operations/linkeddatahub/content/generate_ontology_views.py @@ -1,7 +1,8 @@ +from typing import ClassVar import hashlib from rdflib import URIRef, Literal, Namespace, Graph from rdflib.namespace import RDF, RDFS, XSD, DCTERMS -from web_algebra.operation import Operation +from web_algebra.operation import Operation, OperationKind class GenerateOntologyViews(Operation): @@ -16,6 +17,9 @@ class GenerateOntologyViews(Operation): they yield at most one value, so a table view would be redundant. """ + # composes writes internally rather than through nested forms, so it declares its own kind + kind: ClassVar[OperationKind] = OperationKind.WRITE + @classmethod def name(cls): return "ldh-GenerateOntologyViews" @@ -164,7 +168,7 @@ def _generate_sparql_query(self, property_uri: URIRef) -> str: return sparql - def execute_json(self, arguments: dict, variable_stack: list = []) -> Graph: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Graph: """JSON execution: process arguments with type checking""" # Process ontology graph — accept a Graph (e.g. from CONSTRUCT) or a # JSON-LD document, converting the latter to a Graph diff --git a/src/web_algebra/operations/linkeddatahub/content/generate_portal.py b/src/web_algebra/operations/linkeddatahub/content/generate_portal.py index ceb9507..1cd5461 100644 --- a/src/web_algebra/operations/linkeddatahub/content/generate_portal.py +++ b/src/web_algebra/operations/linkeddatahub/content/generate_portal.py @@ -1,7 +1,8 @@ +from typing import ClassVar from rdflib import URIRef, Literal from rdflib.namespace import XSD from rdflib.query import Result -from web_algebra.operation import Operation +from web_algebra.operation import Operation, OperationKind from web_algebra.operations.schema.extract_ontology import ExtractOntology from web_algebra.operations.linkeddatahub.content.generate_ontology_views import GenerateOntologyViews from web_algebra.operations.linkeddatahub.content.generate_class_containers import GenerateClassContainers @@ -20,6 +21,9 @@ class GeneratePortal(Operation): 4. GenerateClassContainers - creates containers for each class with instance views """ + # composes writes internally rather than through nested forms, so it declares its own kind + kind: ClassVar[OperationKind] = OperationKind.WRITE + @classmethod def name(cls): return "ldh-GeneratePortal" @@ -115,7 +119,7 @@ def execute(self, endpoint: URIRef, ontology_namespace: URIRef, parent_container return JSONResult(list(all_vars), all_bindings) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Result: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Result: """JSON execution: process arguments with strict type checking""" # Process endpoint endpoint_data = Operation.process_json( diff --git a/src/web_algebra/operations/linkeddatahub/content/remove_block.py b/src/web_algebra/operations/linkeddatahub/content/remove_block.py index 241f51b..b37fd7a 100644 --- a/src/web_algebra/operations/linkeddatahub/content/remove_block.py +++ b/src/web_algebra/operations/linkeddatahub/content/remove_block.py @@ -40,7 +40,7 @@ def inputSchema(cls) -> dict: "required": ["url"], } - def execute_json(self, arguments: dict[str, str], variable_stack: list = []) -> Any: + def execute_json(self, arguments: dict[str, str], variable_stack: list = None) -> Any: """JSON execution: process arguments and delegate to execute()""" # Process required arguments url_data = Operation.process_json( diff --git a/src/web_algebra/operations/linkeddatahub/create_container.py b/src/web_algebra/operations/linkeddatahub/create_container.py index 59ff34b..7ac40db 100644 --- a/src/web_algebra/operations/linkeddatahub/create_container.py +++ b/src/web_algebra/operations/linkeddatahub/create_container.py @@ -113,7 +113,7 @@ def execute( # Call parent PUT execute method return super().execute(URIRef(url), graph) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Result: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Result: """JSON execution: process arguments and call pure function""" # Process parent URI parent_data = Operation.process_json( diff --git a/src/web_algebra/operations/linkeddatahub/create_item.py b/src/web_algebra/operations/linkeddatahub/create_item.py index 1065e9f..cf5c27e 100644 --- a/src/web_algebra/operations/linkeddatahub/create_item.py +++ b/src/web_algebra/operations/linkeddatahub/create_item.py @@ -96,7 +96,7 @@ def execute( # Call parent PUT execute method return super().execute(URIRef(url), graph) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Result: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Result: """JSON execution: process arguments with strict type checking""" # Process container URI container_data = Operation.process_json( diff --git a/src/web_algebra/operations/linkeddatahub/list.py b/src/web_algebra/operations/linkeddatahub/list.py index 67e394c..f1971ce 100644 --- a/src/web_algebra/operations/linkeddatahub/list.py +++ b/src/web_algebra/operations/linkeddatahub/list.py @@ -102,7 +102,7 @@ def execute( return result def execute_json( - self, arguments: dict[str, str], variable_stack: list = [] + self, arguments: dict[str, str], variable_stack: list = None ) -> list[dict]: """JSON execution: process arguments and delegate to execute()""" # Process required arguments diff --git a/src/web_algebra/operations/merge.py b/src/web_algebra/operations/merge.py index 278cfe8..749f8cc 100644 --- a/src/web_algebra/operations/merge.py +++ b/src/web_algebra/operations/merge.py @@ -54,7 +54,7 @@ def execute(self, graphs: List[Graph]) -> Graph: logging.info("Merged RDF data (%s triple(s))", len(merged_graph)) return merged_graph - def execute_json(self, arguments: dict, variable_stack: list = []) -> Graph: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Graph: """JSON execution: process arguments and delegate to execute()""" # Process graphs argument - may return dicts with processed nested operations graphs_data = Operation.process_json( diff --git a/src/web_algebra/operations/position.py b/src/web_algebra/operations/position.py new file mode 100644 index 0000000..3d00d74 --- /dev/null +++ b/src/web_algebra/operations/position.py @@ -0,0 +1,36 @@ +from rdflib import Literal +from rdflib.namespace import XSD + +from web_algebra.exceptions import NoFocusError +from web_algebra.focus import Focus +from web_algebra.operation import Operation + + +class Position(Operation): + """ + Returns the 1-based position of the current focus item, per XPath's + fn:position(). + """ + + @classmethod + def description(cls) -> str: + return """Returns the 1-based position of the current iteration item, per XPath's fn:position(). + + Only meaningful inside ForEach, which establishes the focus + (item, position, size); within a focus, 1 <= position <= size.""" + + @classmethod + def inputSchema(cls) -> dict: + return {"type": "object", "properties": {}, "additionalProperties": False} + + def execute(self, focus: Focus) -> Literal: + """Pure function: focus → 1-based position as xsd:integer""" + if not isinstance(focus, Focus): + raise NoFocusError( + "Position requires an iteration focus (only ForEach establishes one)" + ) + return Literal(focus.position, datatype=XSD.integer) + + def execute_json(self, arguments: dict, variable_stack: list = None) -> Literal: + """JSON execution: read the position from the current focus""" + return self.execute(self.context) diff --git a/src/web_algebra/operations/resolve_uri.py b/src/web_algebra/operations/resolve_uri.py index fdf1d34..3873fca 100644 --- a/src/web_algebra/operations/resolve_uri.py +++ b/src/web_algebra/operations/resolve_uri.py @@ -51,7 +51,7 @@ def execute(self, base: URIRef, relative: Literal) -> URIRef: resolved_uri = urljoin(base_str, relative_str) return URIRef(resolved_uri) - def execute_json(self, arguments: dict, variable_stack: list = []) -> URIRef: + def execute_json(self, arguments: dict, variable_stack: list = None) -> URIRef: """JSON execution: process arguments and call pure function""" # Process base URI base_data = Operation.process_json( diff --git a/src/web_algebra/operations/schema/extract_classes.py b/src/web_algebra/operations/schema/extract_classes.py index 9a68470..db4dae7 100644 --- a/src/web_algebra/operations/schema/extract_classes.py +++ b/src/web_algebra/operations/schema/extract_classes.py @@ -1,25 +1,14 @@ -from rdflib import URIRef, Literal, Graph -from rdflib.namespace import XSD -from web_algebra.operations.sparql.construct import CONSTRUCT +from typing import ClassVar from web_algebra.operation import Operation +from web_algebra.schema_extraction import SchemaExtraction -class ExtractClasses(CONSTRUCT): - @classmethod - def description(cls) -> str: - return "Extracts OWL classes from an RDF dataset." - - @classmethod - def inputSchema(cls) -> dict: - return { - "type": "object", - "properties": {"endpoint": {"type": "string"}}, - "required": ["endpoint"], - } +class ExtractClasses(SchemaExtraction, Operation): + """formal-semantics.md §4.6. The query is REST-VKG's, so the two + serializations extract the same schema; `%SCOPE%` is where the + `bindings` VALUES block goes.""" - def execute(self, endpoint: URIRef) -> Graph: - """Pure function: extract OWL classes with RDFLib terms""" - query = Literal(""" + QUERY: ClassVar[str] = """ PREFIX owl: PREFIX rdfs: @@ -28,28 +17,19 @@ def execute(self, endpoint: URIRef) -> Graph: ?class a owl:Class . } WHERE - { { ?instance a ?class + { %SCOPE% + { ?subject a ?class FILTER ( ! isBlank(?class) ) } UNION { GRAPH ?g - { ?instance a ?class + { ?subject a ?class FILTER ( ! isBlank(?class) ) } } } -""", datatype=XSD.string) - return super().execute(endpoint, query) +""" - def execute_json(self, arguments: dict, variable_stack: list = []) -> Graph: - """JSON execution: process arguments with strict type checking""" - # Process endpoint - endpoint_data = Operation.process_json( - self.settings, arguments["endpoint"], self.context, variable_stack - ) - if not isinstance(endpoint_data, URIRef): - raise TypeError( - f"ExtractClasses operation expects 'endpoint' to be URIRef, got {type(endpoint_data)}" - ) - - return self.execute(endpoint_data) + @classmethod + def description(cls) -> str: + return "Extracts OWL classes (owl:Class candidates) from rdf:type usage in an RDF dataset, optionally scoped to the subjects in 'bindings'." diff --git a/src/web_algebra/operations/schema/extract_datatype_properties.py b/src/web_algebra/operations/schema/extract_datatype_properties.py index a034b20..8cf7e80 100644 --- a/src/web_algebra/operations/schema/extract_datatype_properties.py +++ b/src/web_algebra/operations/schema/extract_datatype_properties.py @@ -1,119 +1,55 @@ -from rdflib import URIRef, Literal, Graph -from rdflib.namespace import XSD -from web_algebra.operations.sparql.construct import CONSTRUCT +from typing import ClassVar from web_algebra.operation import Operation +from web_algebra.schema_extraction import SchemaExtraction -class ExtractDatatypeProperties(CONSTRUCT): - @classmethod - def description(cls) -> str: - return "Extracts OWL datatype properties from an RDF dataset." - - @classmethod - def inputSchema(cls) -> dict: - return { - "type": "object", - "properties": {"endpoint": {"type": "string"}}, - "required": ["endpoint"], - } - - def execute(self, endpoint: URIRef) -> Graph: - """Pure function: extract OWL datatype properties with RDFLib terms +class ExtractDatatypeProperties(SchemaExtraction, Operation): + """formal-semantics.md §4.6. The query is REST-VKG's, so the two + serializations extract the same schema; `%SCOPE%` is where the + `bindings` VALUES block goes.""" - Infers functional properties using closed world assumption: - - Counts max cardinality by examining all subjects in the dataset - - Creates OWL restriction with maxQualifiedCardinality - - When maxC = 1, property is inferred to be functional in this dataset - - Note: Inference based solely on present data, not formal ontology definitions - """ - query = Literal(""" + QUERY: ClassVar[str] = """ PREFIX rdf: PREFIX rdfs: PREFIX owl: - CONSTRUCT { - # Basic property metadata - constructed for ALL properties - ?property a owl:DatatypeProperty ; - rdfs:domain ?domain ; - rdfs:range ?datatype . - - # Restriction triples - only constructed when ?restriction and ?maxCardinality are bound - # This happens only for functional properties (maxC = 1) + ?property a owl:DatatypeProperty ; rdfs:domain ?domain ; rdfs:range ?datatype . ?domain rdfs:subClassOf ?restriction . - - ?restriction a owl:Restriction ; - owl:onProperty ?property ; - owl:maxQualifiedCardinality ?maxCardinality ; - owl:onDataRange ?datatype . + ?restriction a owl:Restriction ; owl:onProperty ?property ; owl:maxQualifiedCardinality ?maxCardinality ; owl:onDataRange ?datatype . } WHERE { { - # Outermost SELECT: Conditionally bind maxCardinality only when maxC = 1 - # The IF expression makes ?maxCardinality unbound for non-functional properties - SELECT ?domain ?property ?datatype (IF(?maxC = 1, ?maxC, ?UNDEF) AS ?maxCardinality) + SELECT ?property ?datatype (IF(?domains = 1, ?aDomain, ?UNDEF) AS ?domain) (IF(?maxC = 1 && ?domains = 1, 1, ?UNDEF) AS ?maxCardinality) WHERE { { - # Outer SELECT: Aggregate to single domain per property-datatype pair - # Filter to properties with unambiguous domain (COUNT DISTINCT <= 1) - # Calculate max cardinality across all subjects - SELECT ?property ?datatype (SAMPLE(?d) AS ?domain) (MAX(?c) AS ?maxC) + SELECT ?property ?datatype (MAX(?c) AS ?maxC) (COUNT(DISTINCT ?type) AS ?domains) (SAMPLE(?type) AS ?aDomain) WHERE { { - # Inner SELECT: Count literals per subject-property-datatype triple - # This gives us cardinality for each individual subject SELECT ?subject ?property ?datatype (COUNT(?literal) AS ?c) WHERE { - { - ?subject ?property ?literal . - FILTER(?property != rdf:type) - FILTER(isLiteral(?literal)) - BIND(datatype(?literal) AS ?datatype) - } UNION { - GRAPH ?g { - ?subject ?property ?literal . - FILTER(?property != rdf:type) - FILTER(isLiteral(?literal)) - BIND(datatype(?literal) AS ?datatype) - } - } + %SCOPE% + { ?subject ?property ?literal . FILTER(?property != rdf:type) FILTER(isLiteral(?literal)) BIND(datatype(?literal) AS ?datatype) } + UNION + { GRAPH ?g { ?subject ?property ?literal . FILTER(?property != rdf:type) FILTER(isLiteral(?literal)) BIND(datatype(?literal) AS ?datatype) } } } GROUP BY ?subject ?property ?datatype } - OPTIONAL { - { ?subject a ?d } - UNION - { GRAPH ?subjG { ?subject a ?d } } - FILTER(!isBlank(?d)) + SELECT ?subject (SAMPLE(?d) AS ?type) WHERE { + %SCOPE% + { ?subject a ?d } UNION { GRAPH ?subjG { ?subject a ?d } } + FILTER(!isBlank(?d)) + } GROUP BY ?subject } } GROUP BY ?property ?datatype - HAVING(COUNT(DISTINCT ?d) <= 1) } } } - - # Create blank node for restriction ONLY when ?maxCardinality is bound (functional properties) - # This OPTIONAL block does NOT filter out rows - all properties are still returned - # For functional properties: ?restriction is bound, and restriction triples are constructed - # For non-functional properties: ?restriction is unbound, no restriction triples created - OPTIONAL { - FILTER(BOUND(?maxCardinality)) - BIND(BNODE() AS ?restriction) - } + OPTIONAL { FILTER(BOUND(?maxCardinality)) BIND(BNODE() AS ?restriction) } } -""", datatype=XSD.string) - return super().execute(endpoint, query) +""" - def execute_json(self, arguments: dict, variable_stack: list = []) -> Graph: - """JSON execution: process arguments with strict type checking""" - # Process endpoint - endpoint_data = Operation.process_json( - self.settings, arguments["endpoint"], self.context, variable_stack - ) - if not isinstance(endpoint_data, URIRef): - raise TypeError( - f"ExtractDatatypeProperties operation expects 'endpoint' to be URIRef, got {type(endpoint_data)}" - ) - - return self.execute(endpoint_data) + @classmethod + def description(cls) -> str: + return "Extracts OWL datatype properties from literal-valued predicates in an RDF dataset, optionally scoped to the subjects in 'bindings'." diff --git a/src/web_algebra/operations/schema/extract_object_properties.py b/src/web_algebra/operations/schema/extract_object_properties.py index fe5b2b7..53ee6ed 100644 --- a/src/web_algebra/operations/schema/extract_object_properties.py +++ b/src/web_algebra/operations/schema/extract_object_properties.py @@ -1,109 +1,63 @@ -from rdflib import URIRef, Literal, Graph -from rdflib.namespace import XSD -from web_algebra.operations.sparql.construct import CONSTRUCT +from typing import ClassVar from web_algebra.operation import Operation +from web_algebra.schema_extraction import SchemaExtraction -class ExtractObjectProperties(CONSTRUCT): - @classmethod - def description(cls) -> str: - return "Extracts OWL object properties from an RDF dataset." - - @classmethod - def inputSchema(cls) -> dict: - return { - "type": "object", - "properties": {"endpoint": {"type": "string"}}, - "required": ["endpoint"], - } - - def execute(self, endpoint: URIRef) -> Graph: - """Pure function: extract OWL object properties with RDFLib terms +class ExtractObjectProperties(SchemaExtraction, Operation): + """formal-semantics.md §4.6. The query is REST-VKG's, so the two + serializations extract the same schema; `%SCOPE%` is where the + `bindings` VALUES block goes.""" - Infers functional properties using closed world assumption: - - Counts max cardinality per property across all subjects in the dataset - - When global max = 1, emits ?property a owl:FunctionalProperty - - Note: Inference based solely on present data, not formal ontology definitions - """ - query = Literal(""" + QUERY: ClassVar[str] = """ PREFIX rdf: PREFIX rdfs: PREFIX owl: - CONSTRUCT { - ?property a owl:ObjectProperty ; - rdfs:domain ?domain ; - rdfs:range ?range . + ?property a owl:ObjectProperty ; rdfs:domain ?domain ; rdfs:range ?range . ?functional a owl:FunctionalProperty . } WHERE { { - SELECT ?property ?domain ?range (IF(?maxC = 1, ?property, ?UNDEF) AS ?functional) + SELECT ?property ?range (IF(?domains = 1, ?aDomain, ?UNDEF) AS ?domain) (IF(?maxC = 1, ?property, ?UNDEF) AS ?functional) WHERE { { - SELECT ?property (SAMPLE(?d) AS ?domain) (SAMPLE(?r) AS ?range) (MAX(?maxC2) AS ?maxC) + SELECT ?property ?domains ?aDomain ?maxC (SAMPLE(?objectType) AS ?range) WHERE { { - SELECT ?property ?r (SAMPLE(?d2) AS ?d) (MAX(?c) AS ?maxC2) + SELECT ?property (MAX(?c) AS ?maxC) (COUNT(DISTINCT ?type) AS ?domains) (SAMPLE(?type) AS ?aDomain) (SAMPLE(?anObject) AS ?sampleObject) WHERE { { - SELECT ?subject ?property ?r (COUNT(?object) AS ?c) + SELECT ?subject ?property (COUNT(?object) AS ?c) (SAMPLE(?object) AS ?anObject) WHERE { - { - ?subject ?property ?object . - FILTER(?property != rdf:type) - FILTER(!isLiteral(?object)) - OPTIONAL { - { ?object a ?r } - UNION - { GRAPH ?objG { ?object a ?r } } - FILTER(!isBlank(?r)) - } - } UNION { - GRAPH ?g { - ?subject ?property ?object . - FILTER(?property != rdf:type) - FILTER(!isLiteral(?object)) - OPTIONAL { - { ?object a ?r } - UNION - { GRAPH ?objG { ?object a ?r } } - FILTER(!isBlank(?r)) - } - } - } + %SCOPE% + { ?subject ?property ?object . FILTER(?property != rdf:type) FILTER(!isLiteral(?object)) } + UNION + { GRAPH ?g { ?subject ?property ?object . FILTER(?property != rdf:type) FILTER(!isLiteral(?object)) } } } - GROUP BY ?subject ?property ?r + GROUP BY ?subject ?property } OPTIONAL { - { ?subject a ?d2 } - UNION - { GRAPH ?subjG { ?subject a ?d2 } } - FILTER(!isBlank(?d2)) + SELECT ?subject (SAMPLE(?d) AS ?type) WHERE { + %SCOPE% + { ?subject a ?d } UNION { GRAPH ?subjG { ?subject a ?d } } + FILTER(!isBlank(?d)) + } GROUP BY ?subject } } - GROUP BY ?property ?r - HAVING(COUNT(DISTINCT ?d2) <= 1) + GROUP BY ?property + } + OPTIONAL { + { ?sampleObject a ?objectType } UNION { GRAPH ?objG { ?sampleObject a ?objectType } } + FILTER(!isBlank(?objectType)) } } - GROUP BY ?property - HAVING(COUNT(DISTINCT ?r) <= 1) + GROUP BY ?property ?domains ?aDomain ?maxC } } } } -""", datatype=XSD.string) - return super().execute(endpoint, query) - - def execute_json(self, arguments: dict, variable_stack: list = []) -> Graph: - """JSON execution: process arguments with strict type checking""" - # Process endpoint - endpoint_data = Operation.process_json( - self.settings, arguments["endpoint"], self.context, variable_stack - ) - if not isinstance(endpoint_data, URIRef): - raise TypeError( - f"ExtractObjectProperties operation expects 'endpoint' to be URIRef, got {type(endpoint_data)}" - ) +""" - return self.execute(endpoint_data) + @classmethod + def description(cls) -> str: + return "Extracts OWL object properties (and functional properties, closed-world) from IRI-valued predicates in an RDF dataset, optionally scoped to the subjects in 'bindings'." diff --git a/src/web_algebra/operations/schema/extract_ontology.py b/src/web_algebra/operations/schema/extract_ontology.py index 055fff7..5f272f4 100644 --- a/src/web_algebra/operations/schema/extract_ontology.py +++ b/src/web_algebra/operations/schema/extract_ontology.py @@ -1,61 +1,27 @@ -from rdflib import URIRef, Graph +from typing import Optional +from rdflib import Graph, URIRef +from rdflib.query import Result from web_algebra.operation import Operation +from web_algebra.schema_extraction import SchemaExtraction from web_algebra.operations.schema.extract_classes import ExtractClasses from web_algebra.operations.schema.extract_datatype_properties import ExtractDatatypeProperties from web_algebra.operations.schema.extract_object_properties import ExtractObjectProperties -from web_algebra.operations.merge import Merge -class ExtractOntology(Operation): - """Extracts complete OWL ontology (classes + properties) from an RDF dataset. - - Composes ExtractClasses, ExtractDatatypeProperties, and ExtractObjectProperties, - then merges their results using the Merge operation. +class ExtractOntology(SchemaExtraction, Operation): + """The union of the three schema extractions — classes plus datatype and + object properties — as one graph, each scoped by the same `bindings` + (formal-semantics.md §4.6). """ @classmethod def description(cls) -> str: - return "Extracts complete OWL ontology (classes, datatype properties, and object properties with functional property restrictions) from an RDF dataset." - - @classmethod - def inputSchema(cls) -> dict: - return { - "type": "object", - "properties": {"endpoint": {"type": "string"}}, - "required": ["endpoint"], - } - - def execute(self, endpoint: URIRef) -> Graph: - """Extract complete ontology by composing individual extraction operations""" - if not isinstance(endpoint, URIRef): - raise TypeError( - f"ExtractOntology operation expects 'endpoint' to be URIRef, got {type(endpoint)}" - ) - - # Extract classes - classes_graph = ExtractClasses(settings=self.settings, context=self.context).execute(endpoint) - - # Extract datatype properties (with functional property restrictions) - datatype_props_graph = ExtractDatatypeProperties(settings=self.settings, context=self.context).execute(endpoint) - - # Extract object properties (with functional property restrictions) - object_props_graph = ExtractObjectProperties(settings=self.settings, context=self.context).execute(endpoint) - - # Merge all graphs using the Merge operation - graphs = [classes_graph, datatype_props_graph, object_props_graph] - ontology_graph = Merge(settings=self.settings, context=self.context).execute(graphs) - - return ontology_graph - - def execute_json(self, arguments: dict, variable_stack: list = []) -> Graph: - """JSON execution: process arguments with strict type checking""" - # Process endpoint - endpoint_data = Operation.process_json( - self.settings, arguments["endpoint"], self.context, variable_stack - ) - if not isinstance(endpoint_data, URIRef): - raise TypeError( - f"ExtractOntology operation expects 'endpoint' to be URIRef, got {type(endpoint_data)}" - ) - - return self.execute(endpoint_data) + return "Extracts a complete OWL ontology (classes, datatype properties, and object properties with functional property restrictions) from an RDF dataset, optionally scoped to the subjects in 'bindings'." + + def execute(self, endpoint: URIRef, bindings: Optional[Result] = None) -> Graph: + """Extract the ontology by running every extraction with one scope""" + scope = self.scope(endpoint, bindings) + ontology = Graph() + for extraction in (ExtractClasses, ExtractDatatypeProperties, ExtractObjectProperties): + ontology += extraction(settings=self.settings, context=self.context).extract(endpoint, scope) + return ontology diff --git a/src/web_algebra/operations/sparql/construct.py b/src/web_algebra/operations/sparql/construct.py index a54dbab..d8fa706 100644 --- a/src/web_algebra/operations/sparql/construct.py +++ b/src/web_algebra/operations/sparql/construct.py @@ -1,54 +1,51 @@ import logging -from typing import Any +from typing import Any, ClassVar, Type, Union from rdflib import URIRef, Literal, Graph from rdflib.namespace import XSD from mcp import types from web_algebra.mcp_tool import MCPTool +from web_algebra.client_operation import ClientOperation from web_algebra.operation import Operation from web_algebra.client import SPARQLClient +from web_algebra.query_source import QuerySource -class CONSTRUCT(Operation, MCPTool): +class CONSTRUCT(QuerySource, ClientOperation, Operation, MCPTool): """ Executes a SPARQL CONSTRUCT query against a specified endpoint. """ - def model_post_init(self, __context: Any) -> None: - self.client = SPARQLClient( - cert_pem_path=getattr(self.settings, "cert_pem_path", None), - cert_password=getattr(self.settings, "cert_password", None), - verify_ssl=False, # Optionally disable SSL verification - ) + client_class: ClassVar[Type] = SPARQLClient + @classmethod def description(cls) -> str: - return "Executes a SPARQL CONSTRUCT query." + return "Executes a SPARQL CONSTRUCT query over an endpoint or over a graph in hand." @classmethod def inputSchema(cls) -> dict: return { "type": "object", "properties": { - "endpoint": {"type": "string"}, + "endpoint": {"type": "string", "description": "SPARQL endpoint URL (or give 'graph')"}, + "graph": {"description": "A graph to query locally (or give 'endpoint')"}, "query": {"type": "string"}, }, - "required": ["endpoint", "query"], + "required": ["query"], + "oneOf": [{"required": ["endpoint"]}, {"required": ["graph"]}], } - def execute(self, endpoint: URIRef, query: Literal) -> Graph: - """Pure function: execute SPARQL CONSTRUCT query""" - if not isinstance(endpoint, URIRef): - raise TypeError( - f"CONSTRUCT operation expects endpoint to be URIRef, got {type(endpoint)}" - ) - if not isinstance(query, Literal) or query.datatype != XSD.string: - raise TypeError( - f"CONSTRUCT operation expects query to be string Literal, got {type(query)}" - ) - - endpoint_url = str(endpoint) + def execute(self, source: Union[URIRef, Graph], query: Literal) -> Graph: + """Pure function: execute a SPARQL CONSTRUCT query over `source`, an + endpoint URI or a Graph (formal-semantics.md §4.3)""" + self.check_source(source, query) query_str = str(query) + if isinstance(source, Graph): + logging.info("Executing SPARQL CONSTRUCT over a graph with query:\n%s", query_str) + return source.query(query_str).graph + + endpoint_url = str(source) logging.info( "Executing SPARQL CONSTRUCT on %s with query:\n%s", endpoint_url, query_str ) @@ -59,28 +56,11 @@ def execute(self, endpoint: URIRef, query: Literal) -> Graph: # Convert JSON-LD response to RDF Graph return self.to_graph(json_ld_response) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Graph: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Graph: """JSON execution: process arguments and return Graph (same as execute)""" - # Process endpoint - endpoint_data = Operation.process_json( - self.settings, arguments["endpoint"], self.context, variable_stack - ) - if not isinstance(endpoint_data, URIRef): - raise TypeError( - f"CONSTRUCT operation expects 'endpoint' to be URIRef, got {type(endpoint_data)}" - ) - - # Process query - query_data = Operation.process_json( - self.settings, arguments["query"], self.context, variable_stack - ) - if not isinstance(query_data, Literal) or query_data.datatype != XSD.string: - raise TypeError( - f"CONSTRUCT operation expects 'query' to be string Literal, got {type(query_data)}" - ) - - # Return Graph directly (same as execute) - serialization only at boundaries - return self.execute(endpoint_data, query_data) + source = self.resolve_source(arguments, variable_stack) + query = self.resolve_query(arguments, variable_stack) + return self.execute(source, query) def mcp_run(self, arguments: dict, context: Any = None) -> Any: """MCP execution: plain args → plain results""" diff --git a/src/web_algebra/operations/sparql/describe.py b/src/web_algebra/operations/sparql/describe.py index be82c9a..dfc6ec5 100644 --- a/src/web_algebra/operations/sparql/describe.py +++ b/src/web_algebra/operations/sparql/describe.py @@ -1,54 +1,51 @@ import logging -from typing import Any +from typing import Any, ClassVar, Type, Union from rdflib import URIRef, Literal, Graph from rdflib.namespace import XSD from mcp import types from web_algebra.mcp_tool import MCPTool +from web_algebra.client_operation import ClientOperation from web_algebra.operation import Operation from web_algebra.client import SPARQLClient +from web_algebra.query_source import QuerySource -class DESCRIBE(Operation, MCPTool): +class DESCRIBE(QuerySource, ClientOperation, Operation, MCPTool): """ Executes a SPARQL DESCRIBE query against a specified endpoint and returns a JSON-LD response. """ - def model_post_init(self, __context: Any) -> None: - self.client = SPARQLClient( - cert_pem_path=getattr(self.settings, "cert_pem_path", None), - cert_password=getattr(self.settings, "cert_password", None), - verify_ssl=False, # Optionally disable SSL verification - ) + client_class: ClassVar[Type] = SPARQLClient + @classmethod def description(cls) -> str: - return "Executes a SPARQL DESCRIBE query." + return "Executes a SPARQL DESCRIBE query over an endpoint or over a graph in hand." @classmethod def inputSchema(cls) -> dict: return { "type": "object", "properties": { - "endpoint": {"type": "string"}, + "endpoint": {"type": "string", "description": "SPARQL endpoint URL (or give 'graph')"}, + "graph": {"description": "A graph to query locally (or give 'endpoint')"}, "query": {"type": "string"}, }, - "required": ["endpoint", "query"], + "required": ["query"], + "oneOf": [{"required": ["endpoint"]}, {"required": ["graph"]}], } - def execute(self, endpoint: URIRef, query: Literal) -> Graph: - """Pure function: execute SPARQL DESCRIBE query""" - if not isinstance(endpoint, URIRef): - raise TypeError( - f"DESCRIBE operation expects endpoint to be URIRef, got {type(endpoint)}" - ) - if not isinstance(query, Literal) or query.datatype != XSD.string: - raise TypeError( - f"DESCRIBE operation expects query to be string Literal, got {type(query)}" - ) - - endpoint_url = str(endpoint) + def execute(self, source: Union[URIRef, Graph], query: Literal) -> Graph: + """Pure function: execute a SPARQL DESCRIBE query over `source`, an + endpoint URI or a Graph (formal-semantics.md §4.3)""" + self.check_source(source, query) query_str = str(query) + if isinstance(source, Graph): + logging.info("Executing SPARQL DESCRIBE over a graph with query:\n%s", query_str) + return source.query(query_str).graph + + endpoint_url = str(source) logging.info( "Executing SPARQL DESCRIBE on %s with query:\n%s", endpoint_url, query_str ) @@ -59,28 +56,11 @@ def execute(self, endpoint: URIRef, query: Literal) -> Graph: # Convert JSON-LD response to RDF Graph return self.to_graph(json_ld_response) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Graph: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Graph: """JSON execution: process arguments and return Graph (same as execute)""" - # Process endpoint - endpoint_data = Operation.process_json( - self.settings, arguments["endpoint"], self.context, variable_stack - ) - if not isinstance(endpoint_data, URIRef): - raise TypeError( - f"DESCRIBE operation expects 'endpoint' to be URIRef, got {type(endpoint_data)}" - ) - - # Process query - query_data = Operation.process_json( - self.settings, arguments["query"], self.context, variable_stack - ) - if not isinstance(query_data, Literal) or query_data.datatype != XSD.string: - raise TypeError( - f"DESCRIBE operation expects 'query' to be string Literal, got {type(query_data)}" - ) - - # Return Graph directly (same as execute) - serialization only at boundaries - return self.execute(endpoint_data, query_data) + source = self.resolve_source(arguments, variable_stack) + query = self.resolve_query(arguments, variable_stack) + return self.execute(source, query) def mcp_run(self, arguments: dict, context: Any = None) -> Any: """MCP execution: plain args → plain results""" diff --git a/src/web_algebra/operations/sparql/select.py b/src/web_algebra/operations/sparql/select.py index 2bf9550..a11472a 100644 --- a/src/web_algebra/operations/sparql/select.py +++ b/src/web_algebra/operations/sparql/select.py @@ -1,56 +1,61 @@ -from typing import Any +from typing import Any, ClassVar, Type, Union import logging -from rdflib import URIRef, Literal +from rdflib import Graph, URIRef, Literal from rdflib.namespace import XSD from rdflib.query import Result from mcp import types from web_algebra.mcp_tool import MCPTool +from web_algebra.client_operation import ClientOperation from web_algebra.operation import Operation from web_algebra.client import SPARQLClient +from web_algebra.json_result import JSONResult +from web_algebra.query_source import QuerySource -class SELECT(Operation, MCPTool): +class SELECT(QuerySource, ClientOperation, Operation, MCPTool): """ - Executes SPARQL SELECT queries against endpoints + Executes SPARQL SELECT queries over an endpoint or a graph """ - def model_post_init(self, __context: Any) -> None: - self.client = SPARQLClient( - cert_pem_path=getattr(self.settings, "cert_pem_path", None), - cert_password=getattr(self.settings, "cert_password", None), - verify_ssl=False, - ) + client_class: ClassVar[Type] = SPARQLClient + @classmethod def description(cls) -> str: - return "Executes SPARQL SELECT queries against endpoints" + return "Executes a SPARQL SELECT query over an endpoint or over a graph in hand" @classmethod def inputSchema(cls) -> dict: return { "type": "object", "properties": { - "endpoint": {"type": "string", "description": "SPARQL endpoint URL"}, + "endpoint": {"type": "string", "description": "SPARQL endpoint URL (or give 'graph')"}, + "graph": {"description": "A graph to query locally (or give 'endpoint')"}, "query": {"type": "string", "description": "SPARQL SELECT query"}, }, - "required": ["endpoint", "query"], + "required": ["query"], + "oneOf": [{"required": ["endpoint"]}, {"required": ["graph"]}], } - def execute(self, endpoint: URIRef, query: Literal) -> Result: - """Pure function: execute SPARQL query""" + def execute(self, source: Union[URIRef, Graph], query: Literal) -> Result: + """Pure function: execute a SPARQL SELECT query over `source`, an + endpoint URI or a Graph (formal-semantics.md §4.3)""" # Strict Type Checking before any network side effect. - if not isinstance(endpoint, URIRef): - raise TypeError( - f"SELECT expects endpoint to be URIRef, got {type(endpoint).__name__}" - ) - if not isinstance(query, Literal): - raise TypeError( - f"SELECT expects query to be Literal, got {type(query).__name__}" - ) - - endpoint_url = str(endpoint) + self.check_source(source, query) query_str = str(query) + if isinstance(source, Graph): + logging.info("Executing SPARQL SELECT over a graph with query:\n%s", query_str) + result = source.query(query_str) + return JSONResult( + vars=[str(var) for var in result.vars], + bindings=[ + {str(var): term for var, term in row.items() if term is not None} + for row in result.bindings + ], + ) + + endpoint_url = str(source) logging.info( "Executing SPARQL SELECT on %s with query:\n%s", endpoint_url, query_str ) @@ -62,32 +67,13 @@ def execute(self, endpoint: URIRef, query: Literal) -> Result: len(sparql_json.get("results", {}).get("bindings", [])), ) - # Convert to JSONResult for compatibility - from web_algebra.json_result import JSONResult - return JSONResult.from_json(sparql_json) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Result: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Result: """JSON execution: process arguments with strict type checking""" - # Process endpoint - endpoint_data = Operation.process_json( - self.settings, arguments["endpoint"], self.context, variable_stack - ) - if not isinstance(endpoint_data, URIRef): - raise TypeError( - f"SELECT operation expects 'endpoint' to be URIRef, got {type(endpoint_data)}" - ) - - # Process query - query_data = Operation.process_json( - self.settings, arguments["query"], self.context, variable_stack - ) - if not isinstance(query_data, Literal) or query_data.datatype != XSD.string: - raise TypeError( - f"SELECT operation expects 'query' to be string Literal, got {type(query_data)}" - ) - - return self.execute(endpoint_data, query_data) + source = self.resolve_source(arguments, variable_stack) + query = self.resolve_query(arguments, variable_stack) + return self.execute(source, query) def mcp_run(self, arguments: dict, context: Any = None) -> Any: """MCP execution: plain args → plain results""" diff --git a/src/web_algebra/operations/sparql/substitute.py b/src/web_algebra/operations/sparql/substitute.py index 0141c39..55edea7 100644 --- a/src/web_algebra/operations/sparql/substitute.py +++ b/src/web_algebra/operations/sparql/substitute.py @@ -70,9 +70,12 @@ def execute(self, query: Literal, var: Literal, binding_value: Any) -> Literal: raise TypeError( f"Substitute.execute expects var to be Literal, got {type(var)}" ) - if not isinstance(binding_value, (URIRef, Literal, BNode)): + if not isinstance(binding_value, (URIRef, Literal)): + # formal-semantics.md §4.3: a blank-node label in a query is a + # fresh variable, not a reference — substituting one is + # meaningless, so BNode is rejected along with non-Terms. raise TypeError( - f"Substitute.execute expects binding_value to be URIRef, Literal, or BNode, got {type(binding_value)}" + f"Substitute.execute expects binding_value to be URIRef or Literal, got {type(binding_value)}" ) query_str = str(query) @@ -86,7 +89,7 @@ def execute(self, query: Literal, var: Literal, binding_value: Any) -> Literal: return Literal(substituted_query, datatype=XSD.string) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Literal: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Literal: """JSON execution: process arguments and call pure function""" # Process query query_data = Operation.process_json( @@ -185,8 +188,6 @@ def _format_node(self, node): return f'"{node}"^^<{node.datatype}>' else: return f'"{node}"' - elif isinstance(node, BNode): - return f"_: {node}" else: raise ValueError("Unsupported RDFLib node type") diff --git a/src/web_algebra/operations/sparql/values.py b/src/web_algebra/operations/sparql/values.py index f1f9cb6..3f61d12 100644 --- a/src/web_algebra/operations/sparql/values.py +++ b/src/web_algebra/operations/sparql/values.py @@ -81,19 +81,21 @@ def execute( for binding in (data.bindings or []) ] - block = self._render_values(columns, rows) + block = self.render_values(columns, rows) return Literal(f"{str(query)} {block}", datatype=XSD.string) - def _render_values(self, columns: List[str], rows: List[dict]) -> str: - """Render a SPARQL VALUES block from column names and normalised rows.""" + @classmethod + def render_values(cls, columns: List[str], rows: List[dict]) -> str: + """Render a SPARQL VALUES block from column names and normalised rows. + Shared with the schema operations' `bindings` scope (§4.6).""" if len(columns) == 1: col = columns[0] - cells = " ".join(self._format_term(row.get(col)) for row in rows) + cells = " ".join(cls._format_term(row.get(col)) for row in rows) return f"VALUES ?{col} {{ {cells} }}" header = " ".join(f"?{col}" for col in columns) tuples = " ".join( - "( " + " ".join(self._format_term(row.get(col)) for col in columns) + " )" + "( " + " ".join(cls._format_term(row.get(col)) for col in columns) + " )" for row in rows ) return f"VALUES ({header}) {{ {tuples} }}" @@ -115,7 +117,7 @@ def _format_term(term: Optional[Node]) -> str: # "lex"^^
, and bare numeric/boolean forms. return term.n3() - def execute_json(self, arguments: dict, variable_stack: list = []) -> Literal: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Literal: """JSON execution: process arguments and delegate to execute().""" query_data = Operation.process_json( self.settings, arguments["query"], self.context, variable_stack diff --git a/src/web_algebra/operations/sparql_string.py b/src/web_algebra/operations/sparql_string.py index ebf799a..a327765 100644 --- a/src/web_algebra/operations/sparql_string.py +++ b/src/web_algebra/operations/sparql_string.py @@ -1,28 +1,77 @@ import logging -from typing import Any -from rdflib import Literal +import re +import urllib.error +import urllib.parse +import urllib.request +from dataclasses import dataclass +from datetime import datetime +from typing import Any, ClassVar, List, Optional, Type +from rdflib import Graph, Literal, URIRef from rdflib.namespace import XSD +from rdflib.plugins.sparql import prepareQuery +from rdflib.query import Result from mcp import types from openai import OpenAI +from web_algebra.client import SPARQLClient +from web_algebra.client_operation import ClientOperation from web_algebra.mcp_tool import MCPTool from web_algebra.operation import Operation +SYSTEM_PROMPT = ( + "You are a SPARQL expert. Convert the user's natural language question into a valid SPARQL query. " + "When a section titled 'What the endpoint holds' follows, it is the ground truth about the data: use its terms and nothing else. " + "Always include all necessary PREFIX declarations at the beginning of the query. " + "Return only the raw SPARQL query string with no markdown fences, explanations, or extra text. " + "SPARQL is evaluated inside-out: a nested SELECT subquery is evaluated in isolation before its results are joined with the outer WHERE clause. " + "Therefore, any variable used in a FILTER or ORDER BY inside a nested SELECT MUST also be bound by a triple pattern inside that same nested SELECT." +) -class SPARQLString(Operation, MCPTool): +# The prologue of a query: its BASE and PREFIX declarations. +_PROLOGUE = re.compile( + r"^\s*(?:(?:BASE\s*<[^>]*>|PREFIX\s+[^\s:]*:\s*<[^>]*>)\s*)*", re.IGNORECASE +) +_FENCE = re.compile(r"^\s*```[a-zA-Z]*\s*\n(.*?)\n\s*```\s*$", re.DOTALL) + + +@dataclass(frozen=True) +class Rejection: + """Why a query goes back to the model. Fatal when the query would fail + downstream too (it is not a query, not the declared shape, or refused by + the endpoint); not when it merely matches nothing, since an empty answer + can be the true one.""" + + reason: str + fatal: bool + + +class SPARQLString(ClientOperation, Operation, MCPTool): """ - Converts a natural language question into a SPARQL query using OpenAI API. + Writes a SPARQL query for an endpoint from a natural-language question via + an LLM — the algebra's `xsl:evaluate` (formal-semantics.md §4.3). """ + client_class: ClassVar[Type] = SPARQLClient + + #: How many times the model is asked for one query: the first answer and + #: the corrections of it. + MAX_ATTEMPTS: ClassVar[int] = 3 + #: What a context value is cut to in the prompt, in rows and characters. + MAX_CONTEXT_ROWS: ClassVar[int] = 60 + MAX_CONTEXT_CHARS: ClassVar[int] = 100_000 + def model_post_init(self, __context: Any) -> None: - self.client = OpenAI(api_key=getattr(self.settings, "openai_api_key", None)) + super().model_post_init(__context) + self.llm = OpenAI(api_key=getattr(self.settings, "openai_api_key", None)) self.model = getattr(self.settings, "openai_model", None) @classmethod def description(cls) -> str: return """ - Converts a natural language question into a SPARQL query string. - This operation uses OpenAI's API to generate a structured SPARQL query based on the provided question. - The generated query will include necessary PREFIX declarations and will be formatted as a valid SPARQL query string. + Writes a SPARQL query for an endpoint from a natural language question, using an LLM. + `projection` names the variables the query must project (because what follows reads them from its rows). + `context` holds operations (e.g. a SELECT that inventories predicates, or looks a label up) whose + results are shown to the model as what the endpoint holds, before it writes the query. + A query that does not parse or lacks the projection goes back to the model, a few times at most. """ @classmethod @@ -30,56 +79,300 @@ def inputSchema(cls) -> dict: return { "type": "object", "properties": { + "endpoint": { + "type": "string", + "description": "The SPARQL endpoint the query is for.", + }, "question": { "type": "string", "description": "The natural language question to convert into a SPARQL query.", - } + }, + "projection": { + "type": "array", + "items": {"type": "string"}, + "description": "Variable names the query must be a SELECT projecting.", + }, + "context": { + "type": "array", + "description": "Operations whose results show the model what the endpoint holds.", + }, }, - "required": ["question"], + "required": ["endpoint", "question"], } - def execute(self, question: Literal) -> Literal: - """Pure function: generate SPARQL query from question with RDFLib terms""" - if not isinstance(question, Literal): + def execute( + self, + endpoint: URIRef, + question: Literal, + projection: Optional[List[Literal]] = None, + context_values: Optional[List[Any]] = None, + ) -> Literal: + """Generate a query for `endpoint`; `context_values` is the evaluated + JSON `context` argument (renamed: `context` is the operation's focus).""" + if not isinstance(endpoint, URIRef): raise TypeError( - f"SPARQLString.execute expects question to be Literal, got {type(question)}" + f"SPARQLString expects endpoint to be URIRef, got {type(endpoint).__name__}" ) + question = self.to_string_literal(question) + variables = self._variables(projection) + excerpt = self._context_section(context_values or []) - question_str = str(question) - logging.info("Generating SPARQL query for question: %s", question_str) + messages = [ + {"role": "system", "content": self._system_prompt(str(endpoint), variables) + excerpt}, + {"role": "user", "content": str(question)}, + ] + logging.info("Generating SPARQL query for question: %s", question) - # Call OpenAI API to generate SPARQL query - chat_completion = self.client.chat.completions.create( - model=self.model, - messages=[ - { - "role": "system", - "content": "You are an expert in RDF and SPARQL. Generate a valid SPARQL query based on the given natural language question.", - }, - { - "role": "user", - "content": f"Convert this into a SPARQL query:\nQuestion: {question_str}\nRemember to include the necessary PREFIX declarations. Provide only the query string, no explanations or comments or markdown formatting.", - }, - ], - ) + query = "" + rejection: Optional[Rejection] = None + for attempt in range(1, self.MAX_ATTEMPTS + 1): + query = self._strip_fences(self._complete(messages)) + logging.info( + "SPARQLString produced query (attempt %d/%d):\n%s", + attempt, self.MAX_ATTEMPTS, query, + ) + reason = self.rejection(query, variables) + rejection = Rejection(reason, True) if reason else self._unmatched(str(endpoint), query) + if rejection is None: + return Literal(query) + if attempt == self.MAX_ATTEMPTS and not rejection.fatal: + # an empty answer can be the true one; SELECT reports it as zero rows + logging.info("SPARQLString: the last query still matches nothing; it goes downstream") + return Literal(query) + logging.info("SPARQLString rejected the query: %s", rejection.reason) + messages.append({"role": "assistant", "content": query}) + messages.append( + {"role": "user", "content": rejection.reason + " Return the corrected query, and nothing else."} + ) - result = chat_completion.choices[0].message.content - logging.info("Generated SPARQL query: %s", result) - return Literal(result, datatype=XSD.string) + raise ValueError( + f'SPARQLString: after {self.MAX_ATTEMPTS} attempts the model wrote no query for "{question}": ' + f"{rejection.reason}\n{query}" + ) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Literal: - """JSON execution: process arguments and delegate to execute()""" - question_data = Operation.process_json( + def execute_json(self, arguments: dict, variable_stack: list = None) -> Literal: + """JSON execution: evaluate every argument eagerly, then generate""" + endpoint = Operation.process_json( + self.settings, arguments["endpoint"], self.context, variable_stack + ) + question = Operation.process_json( self.settings, arguments["question"], self.context, variable_stack ) - # Allow implicit string conversion - question_literal = self.to_string_literal(question_data) + projection = None + if "projection" in arguments: + projection = Operation.process_json( + self.settings, arguments["projection"], self.context, variable_stack + ) + if not isinstance(projection, list): + raise TypeError( + f"SPARQLString expects 'projection' to be an array of names, got {type(projection).__name__}" + ) + context_values = None + if "context" in arguments: + forms = arguments["context"] + if not isinstance(forms, list): + raise TypeError( + f"SPARQLString expects 'context' to be an array of forms, got {type(forms).__name__}" + ) + # evaluated one by one, not as a sequence form: each is shown as + # one exploration, and a Result is one value, not its rows + context_values = [ + Operation.process_json(self.settings, form, self.context, variable_stack) + for form in forms + ] + return self.execute(endpoint, question, projection, context_values) + + @staticmethod + def _variables(projection: Optional[List[Any]]) -> List[str]: + """The declared projection's names, without SPARQL's `?`/`$` sigil.""" + names: List[str] = [] + for name in projection or []: + if not Operation.is_string_literal(name): + raise TypeError( + f"SPARQLString expects 'projection' names to be string Literals, got {name!r}" + ) + text = str(name).strip() + names.append(text[1:] if text[:1] in ("?", "$") else text) + return names + + @staticmethod + def rejection(query: str, variables: List[str]) -> Optional[str]: + """Why `query` is not the query asked for, as the next thing to tell + the model, or None when it is: it must parse, and with a declared + projection be a SELECT projecting every declared variable.""" + if not query.strip(): + return "You wrote no query." + try: + prepared = prepareQuery(query) + except Exception as e: + return f"That is not a SPARQL query: {str(e).strip()}." + if not variables: + return None + declared = " ".join(f"?{v}" for v in variables) + if prepared.algebra.name != "SelectQuery": + return f"The query must be a SELECT projecting {declared}." + projected = {str(var) for var in prepared.algebra.get("PV", [])} + missing = [v for v in variables if v not in projected] + if not missing: + return None + return ( + f"The query must project {declared}, named exactly so; it does not project " + + " ".join(f"?{v}" for v in missing) + + "." + ) + + def _unmatched(self, endpoint: str, query: str) -> Optional[Rejection]: + """Why the endpoint would answer `query` with nothing, or None when it + would not — or cannot say. The query's pattern is put to the endpoint + as an ASK; a refusal (4xx) goes back to the model with the endpoint's + reason, an answer it could not give lets the query through.""" + ask = self._ask(query) + if ask is None: + return None + try: + answer = self.client.query(endpoint, ask, post=True) + except urllib.error.HTTPError as e: + if 400 <= e.code < 500: + return Rejection( + f"The endpoint refused the query's pattern with HTTP {e.code}: {e.reason}. " + "Rewrite the query so that this endpoint accepts it, with the terms under " + "'What the endpoint holds' only.", + True, + ) + logging.info("SPARQLString: %s answered %s to the ASK; whether the query matches is unknown", endpoint, e.code) + return None + except Exception as e: + logging.info("SPARQLString: could not ask %s whether the query matches: %s", endpoint, e) + return None + if not isinstance(answer, dict) or answer.get("boolean") is not False: + return None + return Rejection( + "The endpoint answered that the query's pattern matches nothing, so the query as written " + "returns no rows. Rewrite it with the terms under 'What the endpoint holds' only, and loosen " + "what you assumed about the values: a language tag, a datatype, an exact string or an " + "identifier you did not see there.", + False, + ) + + @staticmethod + def _ask(query: str) -> Optional[str]: + """A SELECT's pattern as an ASK: the SELECT becomes a subquery of an + ASK with the same prologue, which SPARQL 1.1 allows with its solution + modifiers and VALUES. None for other query forms, and for a SELECT + with a dataset clause, which a subquery cannot carry.""" + try: + prepared = prepareQuery(query) + except Exception: + return None + if prepared.algebra.name != "SelectQuery" or prepared.algebra.get("datasetClause"): + return None + prologue = _PROLOGUE.match(query).group(0) + ask = f"{prologue}ASK {{ {query[len(prologue):]} }}" + try: + prepareQuery(ask) + except Exception: + return None + return ask + + def _system_prompt(self, endpoint: str, variables: List[str]) -> str: + prompt = SYSTEM_PROMPT + if variables: + prompt += ( + "\n\nThe query must be a SELECT that projects the variables " + + " ".join(f"?{v}" for v in variables) + + ", named exactly so, since what follows reads them by name from its rows; " + "it may project others as well." + ) + now = datetime.now().astimezone() + prompt += ( + f"\n\nThe current date and time is: {now.isoformat()} (timezone: {now.tzname()})." + " When generating date literals for SPARQL, express them as xsd:dateTime values adjusted to UTC (Z suffix)." + ) + agents_md = self._agents_md(endpoint) + if agents_md: + prompt += "\n\n# Service documentation:\n" + agents_md + return prompt + + def _agents_md(self, endpoint: str) -> Optional[str]: + """The endpoint's `AGENTS.md`, resolved against its URL, if it has one.""" + url = urllib.parse.urljoin(endpoint, "AGENTS.md") + request = urllib.request.Request(url, headers={"Accept": "text/markdown, text/plain"}) + try: + with self.client.opener.open(request) as response: + return response.read().decode("utf-8") + except Exception as e: + logging.debug("SPARQLString: no AGENTS.md at %s: %s", url, e) + return None + + def _context_section(self, values: List[Any]) -> str: + """The explorations' results, rendered for the model. An empty one is + an error (§4.3): what it assumed does not match the endpoint.""" + if not values: + return "" + parts = [ + "\n\n# What the endpoint holds\n\nThe results of queries already run against this endpoint. " + "Write the query with the terms that appear here - classes, predicates, identifiers, value shapes - " + "and not with terms recalled from elsewhere; a term that does not appear here is one the data does not have.\n" + ] + for value in values: + if self._is_empty(value): + raise ValueError( + "SPARQLString: a 'context' exploration returned nothing, so the query cannot be " + "written from it; what it assumed - a class, a property, an identifier - does not " + "match the endpoint" + ) + parts.append("\n" + self._render(value) + "\n") + return "".join(parts) + + @staticmethod + def _is_empty(value: Any) -> bool: + if isinstance(value, Result): + return len(list(value)) == 0 + if isinstance(value, Graph): + return len(value) == 0 + if isinstance(value, list): + return len(value) == 0 + return value is None or not str(value).strip() + + def _render(self, value: Any) -> str: + """A result set as its rows, a graph as Turtle, anything else as its + string — cut to MAX_CONTEXT_ROWS rows and MAX_CONTEXT_CHARS characters.""" + if isinstance(value, Result): + names = [str(var) for var in (value.vars or [])] + rows = list(value) + lines = ["\t".join(names)] + for row in rows[: self.MAX_CONTEXT_ROWS]: + cells = row.asdict() + lines.append("\t".join( + cells[name].n3() if cells.get(name) is not None else "" for name in names + )) + if len(rows) > self.MAX_CONTEXT_ROWS: + lines.append(f"... {len(rows) - self.MAX_CONTEXT_ROWS} more rows") + text = "\n".join(lines) + elif isinstance(value, Graph): + text = value.serialize(format="turtle") + else: + text = str(value) + if len(text) > self.MAX_CONTEXT_CHARS: + text = text[: self.MAX_CONTEXT_CHARS] + "\n... cut\n" + return text + + def _complete(self, messages: List[dict]) -> str: + chat_completion = self.llm.chat.completions.create( + model=self.model, messages=messages + ) + return chat_completion.choices[0].message.content or "" - return self.execute(question_literal) + @staticmethod + def _strip_fences(text: str) -> str: + match = _FENCE.match(text) + return (match.group(1) if match else text).strip() def mcp_run(self, arguments: dict, context: Any = None) -> Any: """MCP execution: plain args → plain results""" + endpoint = URIRef(arguments["endpoint"]) question = Literal(arguments["question"], datatype=XSD.string) + projection = [Literal(name) for name in arguments.get("projection", [])] - result = self.execute(question) + result = self.execute(endpoint, question, projection or None) return [types.TextContent(type="text", text=str(result))] diff --git a/src/web_algebra/operations/str.py b/src/web_algebra/operations/str.py index f810fad..69d0f79 100644 --- a/src/web_algebra/operations/str.py +++ b/src/web_algebra/operations/str.py @@ -1,19 +1,17 @@ -from typing import Any from rdflib.term import Node -from rdflib import BNode, Literal, URIRef -from rdflib.namespace import XSD -from mcp import types +from rdflib import Literal, URIRef from web_algebra.operation import Operation class Str(Operation): """ - Converts any RDF term to a string literal + Returns the lexical form of a Literal or the codepoint representation of + a URI as a simple literal, per SPARQL 1.1 STR(). """ @classmethod def description(cls) -> str: - return "Converts any RDF term to a string literal" + return "Returns the lexical form of a Literal or the string representation of a URI, per SPARQL's STR() function. The language tag, if any, is not carried over." @classmethod def inputSchema(cls) -> dict: @@ -27,26 +25,19 @@ def inputSchema(cls) -> dict: def execute(self, term: Node) -> Literal: """Pure function: RDFLib term → string literal""" - # Strict Type Checking: spec defines Str as Term → Literal where Term = URI + Literal + BNode. - if not isinstance(term, (URIRef, Literal, BNode)): + # SPARQL 1.1 STR() accepts a literal or an IRI; a blank node is a + # type error (formal-semantics.md §4.2). + if not isinstance(term, (URIRef, Literal)): raise TypeError( - f"Str expects a Term (URIRef, Literal, BNode), got {type(term).__name__}" + f"Str expects a URI or Literal (SPARQL STR), got {type(term).__name__}" ) - # Check if already string-compatible - if isinstance(term, Literal): - if term.datatype == XSD.string: - return term # Already xsd:string, return as-is - elif term.language is not None: - return term # rdf:langString (datatype=None, language=xx), return as-is (compatible) - elif term.datatype is None and term.language is None: - # Plain literal without datatype or language - treat as string - return term - # Convert any other term to xsd:string - return Literal(str(term), datatype=XSD.string) + # The lexical form / codepoint representation as a simple literal + # (no datatype, no language tag) — as rdflib's SPARQL engine does. + return Literal(str(term)) def execute_json( - self, arguments: dict, variable_stack: list = [] + self, arguments: dict, variable_stack: list = None ) -> Literal: """JSON execution: processes JSON args, returns RDFLib string literal""" # Process the input argument through the JSON system @@ -62,14 +53,3 @@ def execute_json( # Call pure function return self.execute(input_data) - - def mcp_run(self, arguments: dict, context: Any = None) -> Any: - """MCP execution: plain args → plain results""" - # Convert plain input to RDFLib term - rdflib_term = Operation.plain_to_rdflib(arguments["input"]) - - # Call pure function - result = self.execute(rdflib_term) - - # Convert result to plain string for MCP - return [types.TextContent(type="text", text=str(result))] diff --git a/src/web_algebra/operations/string/concat.py b/src/web_algebra/operations/string/concat.py index 0182b59..607b976 100644 --- a/src/web_algebra/operations/string/concat.py +++ b/src/web_algebra/operations/string/concat.py @@ -8,12 +8,12 @@ class Concat(Operation, MCPTool): """ - Concatenates multiple string values into a single string. + Concatenates string literals, per SPARQL 1.1 CONCAT(). """ @classmethod def description(cls) -> str: - return "Concatenates multiple string values into a single string." + return "Concatenates multiple string values into a single string, per SPARQL's CONCAT() function: if all inputs carry the same language tag the result carries it too, otherwise the result is a simple (xsd:string) literal." @classmethod def inputSchema(cls) -> dict: @@ -33,17 +33,25 @@ def execute(self, inputs: List[Literal]) -> Literal: """Pure function: concatenate literals with RDFLib terms""" if not isinstance(inputs, list): raise TypeError(f"Concat.execute expects inputs to be list, got {type(inputs)}") - - # Convert all inputs to strings and concatenate - result_str = "" + for input_literal in inputs: if not isinstance(input_literal, Literal): raise TypeError(f"Concat.execute expects all inputs to be Literal, got {type(input_literal)}") - result_str += str(input_literal) - return Literal(result_str, datatype=XSD.string) + result_str = "".join(str(input_literal) for input_literal in inputs) + + # SPARQL 1.1 CONCAT() result kind (formal-semantics.md §4.2): all + # inputs typed xsd:string → xsd:string; all inputs carrying the same + # language tag → that tag; anything else → simple literal. + datatypes = {input_literal.datatype for input_literal in inputs} + languages = {input_literal.language for input_literal in inputs} + if inputs and datatypes == {XSD.string}: + return Literal(result_str, datatype=XSD.string) + if inputs and len(languages) == 1 and None not in languages: + return Literal(result_str, lang=next(iter(languages))) + return Literal(result_str) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Literal: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Literal: """JSON execution: process arguments and call pure function""" inputs_data = arguments["inputs"] diff --git a/src/web_algebra/operations/string/encode_for_uri.py b/src/web_algebra/operations/string/encode_for_uri.py index 0e1dba5..1a2bde7 100644 --- a/src/web_algebra/operations/string/encode_for_uri.py +++ b/src/web_algebra/operations/string/encode_for_uri.py @@ -44,9 +44,10 @@ def execute(self, input_str: Literal) -> Literal: encoded_value = quote(input_value, safe="") # No safe characters logging.info("Encoded URI: %s", encoded_value) - return Literal(encoded_value, datatype=XSD.string) + # simple literal per `simple literal ENCODE_FOR_URI(string literal)` + return Literal(encoded_value) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Literal: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Literal: """JSON execution: process arguments and call pure function""" input_data = Operation.process_json( self.settings, arguments["input"], self.context, variable_stack diff --git a/src/web_algebra/operations/string/replace.py b/src/web_algebra/operations/string/replace.py index 8b2eb46..23ebd0e 100644 --- a/src/web_algebra/operations/string/replace.py +++ b/src/web_algebra/operations/string/replace.py @@ -1,4 +1,4 @@ -from typing import Any +from typing import Any, Optional import logging import re from rdflib import Literal @@ -10,14 +10,15 @@ class Replace(Operation, MCPTool): """ - Replaces occurrences of a specified pattern in an input string with a given replacement. - Aligns with SPARQL's REPLACE() function. + Regular-expression replacement, per SPARQL 1.1 REPLACE() / XPath + fn:replace: `string literal REPLACE(string literal arg, simple literal + pattern, simple literal replacement [, simple literal flags])`. """ @classmethod def description(cls) -> str: - return """Replaces occurrences of a specified pattern in an input string with a given replacement. This operation aligns with SPARQL's REPLACE() function, allowing for flexible string manipulation using regular expressions. - + return """Replaces occurrences of a regular-expression pattern in an input string, per SPARQL's REPLACE() / XPath fn:replace. The replacement string may reference capture groups as $1, $2, ...; flags `s`, `m`, `i`, `x`, `q` are supported. The result is a string literal of the same kind (datatype/language tag) as the input. + Note: this function should not be used to build URIs! That should be done using EncodeForURI()/ResolveURI(). """ @@ -35,69 +36,158 @@ def inputSchema(cls) -> dict: }, "pattern": { "type": "string", - "description": "The pattern to be replaced (regular expression)", + "description": "The regular expression to be replaced (XPath fn:replace syntax)", }, "replacement": { "type": "string", - "description": "The replacement value", + "description": "The replacement string; $1, $2, ... reference capture groups", + }, + "flags": { + "type": "string", + "description": "Optional XPath regex flags: any of s, m, i, x, q", }, }, "required": ["input", "pattern", "replacement"], } + # XPath flag → Python re flag (formal-semantics.md §4.2); `q` is handled + # separately (treat pattern and replacement as literal strings). + _FLAG_MAP = { + "i": re.IGNORECASE, + "s": re.DOTALL, + "m": re.MULTILINE, + "x": re.VERBOSE, + } + + @staticmethod + def _is_string_compatible(lit: Any) -> bool: + # SPARQL string literal: xsd:string, rdf:langString, or plain literal + return isinstance(lit, Literal) and ( + lit.datatype == XSD.string + or (lit.datatype is None and lit.language is not None) + or (lit.datatype is None and lit.language is None) + ) + + @staticmethod + def _is_simple(lit: Any) -> bool: + # SPARQL simple literal (xsd:string accepted per RDF 1.1) + return Operation.is_string_literal(lit) + + @staticmethod + def _translate_replacement(replacement: str) -> str: + """Translate an XPath fn:replace replacement string into Python + `re.sub` syntax: `$N` → `\\g`, `\\$` → `$`, `\\\\` → literal + backslash. Any other use of `\\` or `$` is an error (err:FORX0004). + """ + out = [] + i = 0 + n = len(replacement) + while i < n: + ch = replacement[i] + if ch == "\\": + if i + 1 < n and replacement[i + 1] == "\\": + out.append("\\\\") + i += 2 + continue + if i + 1 < n and replacement[i + 1] == "$": + out.append("$") + i += 2 + continue + raise ValueError( + "Replace: invalid escape in replacement string (XPath err:FORX0004)" + ) + if ch == "$": + j = i + 1 + while j < n and replacement[j].isdigit(): + j += 1 + if j == i + 1: + raise ValueError( + "Replace: '$' must be followed by a group number in the replacement string (XPath err:FORX0004)" + ) + out.append(f"\\g<{replacement[i + 1:j]}>") + i = j + continue + out.append(ch) + i += 1 + return "".join(out) + def execute( self, input_str: Literal, pattern: Literal, replacement: Literal, + flags: Optional[Literal] = None, ) -> Literal: """Pure function: replace pattern in string with RDFLib terms""" - - # Following SPARQL semantics: accept both xsd:string and rdf:langString (language-tagged literals) - def is_string_compatible(lit): - return isinstance(lit, Literal) and ( - lit.datatype == XSD.string # xsd:string - or ( - lit.datatype is None - and lit.language is not None - ) # rdf:langString - or ( - lit.datatype is None - and lit.language is None - ) # plain literal - ) - - if not is_string_compatible(input_str): - raise TypeError( - f"Replace operation expects input to be string-compatible Literal, got {type(input_str)} with datatype {getattr(input_str, 'datatype', None)}" - ) - if not is_string_compatible(pattern): - raise TypeError( - f"Replace operation expects pattern to be string-compatible Literal, got {type(pattern)} with datatype {getattr(pattern, 'datatype', None)}" - ) - if not is_string_compatible(replacement): + if not self._is_string_compatible(input_str): raise TypeError( - f"Replace operation expects replacement to be string-compatible Literal, got {type(replacement)} with datatype {getattr(replacement, 'datatype', None)}" + f"Replace expects input to be a string literal, got {type(input_str)} with datatype {getattr(input_str, 'datatype', None)}" ) + # Per the REPLACE signature, pattern/replacement/flags are simple + # literals — a language-tagged value is a type error. + for name, lit in (("pattern", pattern), ("replacement", replacement)): + if not self._is_simple(lit): + raise TypeError( + f"Replace expects {name} to be a simple literal, got {lit!r}" + ) + if flags is not None and not self._is_simple(flags): + raise TypeError(f"Replace expects flags to be a simple literal, got {flags!r}") input_value = str(input_str) pattern_value = str(pattern) replacement_value = str(replacement) + flags_value = str(flags) if flags is not None else "" + + re_flags = 0 + literal_mode = False + for flag_char in flags_value: + if flag_char == "q": + literal_mode = True + elif flag_char in self._FLAG_MAP: + re_flags |= self._FLAG_MAP[flag_char] + else: + raise ValueError( + f"Replace: invalid flag {flag_char!r} (XPath err:FORX0001)" + ) + + if literal_mode: + # q: pattern and replacement are taken literally + pattern_value = re.escape(pattern_value) + replacement_re = replacement_value.replace("\\", "\\\\") + else: + replacement_re = self._translate_replacement(replacement_value) + + try: + compiled = re.compile(pattern_value, re_flags) + except re.error as e: + raise ValueError(f"Replace: invalid regular expression: {e} (XPath err:FORX0002)") from None + if compiled.search(""): + raise ValueError( + "Replace: pattern matches a zero-length string (XPath err:FORX0003)" + ) logging.info( - "Resolving Replace arguments: input=%s, pattern=%s, replacement=%s", + "Resolving Replace arguments: input=%s, pattern=%s, replacement=%s, flags=%s", input_value, pattern_value, replacement_value, + flags_value, ) - formatted_string = re.sub(pattern_value, replacement_value, input_value) + try: + formatted_string = compiled.sub(replacement_re, input_value) + except re.error as e: + raise ValueError(f"Replace: invalid replacement string: {e} (XPath err:FORX0004)") from None logging.info("Formatted result: %s", formatted_string) - return Literal(formatted_string, datatype=XSD.string) + # SPARQL string-function convention: the result is a string literal + # of the same kind as the first argument. + if input_str.language is not None: + return Literal(formatted_string, lang=input_str.language) + return Literal(formatted_string, datatype=input_str.datatype) def execute_json( - self, arguments: dict, variable_stack: list = [] + self, arguments: dict, variable_stack: list = None ) -> Literal: """JSON execution: process arguments with strict type checking""" # Process input - allow implicit string conversion @@ -118,14 +208,28 @@ def execute_json( ) replacement_literal = self.to_string_literal(replacement_data) - return self.execute(input_literal, pattern_literal, replacement_literal) + flags_literal = None + if "flags" in arguments: + flags_data = Operation.process_json( + self.settings, arguments["flags"], self.context, variable_stack + ) + flags_literal = self.to_string_literal(flags_data) + + return self.execute( + input_literal, pattern_literal, replacement_literal, flags_literal + ) def mcp_run(self, arguments: dict, context: Any = None) -> Any: """MCP execution: plain args → plain results""" input_str = Literal(arguments["input"], datatype=XSD.string) pattern = Literal(arguments["pattern"], datatype=XSD.string) replacement = Literal(arguments["replacement"], datatype=XSD.string) + flags = ( + Literal(arguments["flags"], datatype=XSD.string) + if "flags" in arguments + else None + ) - result = self.execute(input_str, pattern, replacement) + result = self.execute(input_str, pattern, replacement, flags) return [types.TextContent(type="text", text=str(result))] diff --git a/src/web_algebra/operations/struuid.py b/src/web_algebra/operations/struuid.py index 5443e18..358133a 100644 --- a/src/web_algebra/operations/struuid.py +++ b/src/web_algebra/operations/struuid.py @@ -2,7 +2,6 @@ import uuid from typing import Any from rdflib import Literal -from rdflib.namespace import XSD from mcp import types from web_algebra.operation import Operation from web_algebra.mcp_tool import MCPTool @@ -27,9 +26,10 @@ def execute(self) -> Literal: generated_uuid = str(uuid.uuid4()) logging.info("Generated UUID: %s", generated_uuid) - return Literal(generated_uuid, datatype=XSD.string) + # simple literal per `simple literal STRUUID()` + return Literal(generated_uuid) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Literal: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Literal: """JSON execution: process arguments and call pure function""" return self.execute() diff --git a/src/web_algebra/operations/uri.py b/src/web_algebra/operations/uri.py index cf94ec2..32b9787 100644 --- a/src/web_algebra/operations/uri.py +++ b/src/web_algebra/operations/uri.py @@ -1,7 +1,5 @@ -from typing import Any -from rdflib import URIRef +from rdflib import BNode, URIRef from rdflib.term import Node -from mcp import types from web_algebra.operation import Operation @@ -30,10 +28,13 @@ def execute(self, term: Node) -> URIRef: raise TypeError( f"URI operation expects input to be RDFLib term, got {type(term)}" ) + if isinstance(term, BNode): + # formal-semantics.md §4.2: a blank node has no IRI to cast to. + raise TypeError("URI cannot cast a BNode — blank nodes have no IRI") return URIRef(str(term)) - def execute_json(self, arguments: dict, variable_stack: list = []) -> URIRef: + def execute_json(self, arguments: dict, variable_stack: list = None) -> URIRef: """JSON execution: processes JSON args, returns RDFLib URI reference""" # Process the input argument through the JSON system input_data = Operation.process_json( @@ -48,14 +49,3 @@ def execute_json(self, arguments: dict, variable_stack: list = []) -> URIRef: # Call pure function return self.execute(input_data) - - def mcp_run(self, arguments: dict, context: Any = None) -> Any: - """MCP execution: plain args → plain results""" - # Convert plain input to RDFLib term - rdflib_term = Operation.plain_to_rdflib(arguments["input"]) - - # Call pure function - result = self.execute(rdflib_term) - - # Convert result to plain string for MCP - return [types.TextContent(type="text", text=str(result))] diff --git a/src/web_algebra/operations/value.py b/src/web_algebra/operations/value.py index b7e0125..85f774e 100644 --- a/src/web_algebra/operations/value.py +++ b/src/web_algebra/operations/value.py @@ -1,7 +1,9 @@ +from collections.abc import Mapping from typing import Any import logging from rdflib.query import ResultRow -from mcp import types +from web_algebra.exceptions import VariableNotFoundError +from web_algebra.focus import Focus from web_algebra.operation import Operation @@ -38,29 +40,36 @@ def execute(self, name: str, context: Any, variable_stack: list) -> Any: ) return result else: - # Context binding reference + # Focus-item lookup (formal-semantics.md §3.5). The item shapes + # are closed: Binding → bound term, mapping → member value; + # anything else is an error. + if isinstance(context, Focus): + context = context.item if isinstance(context, ResultRow): # SPARQL result row - access by variable name try: return context[name] # Already RDFLib term except KeyError: - raise ValueError(f"Variable '{name}' not found in ResultRow") + raise VariableNotFoundError( + f"Variable '{name}' not found in ResultRow" + ) + elif isinstance(context, Mapping): + if name in context: + return context[name] + raise VariableNotFoundError( + f"Context member '{name}' not found in mapping" + ) else: - # Other context types - if hasattr(context, name): - return getattr(context, name) raise ValueError( - f"Context variable '{name}' not found in {type(context)}" + f"Value cannot look up '{name}' in a {type(context).__name__} " + "focus item (expected a Binding or a mapping)" ) - def execute_json(self, arguments: dict, variable_stack: list = []) -> Any: + def execute_json(self, arguments: dict, variable_stack: list = None) -> Any: """JSON execution: processes JSON args, returns value (RDFLib term or raw value)""" + if variable_stack is None: + variable_stack = [] var_name: str = arguments["name"] logging.info("Resolving Value variable: %s", var_name) return self.execute(var_name, self.context, variable_stack) - - def mcp_run(self, arguments: dict, context: Any = None) -> Any: - """MCP execution: plain args → plain results""" - # For MCP, we don't have variable stack or context, so just return the name - return [types.TextContent(type="text", text=arguments["name"])] diff --git a/src/web_algebra/operations/variable.py b/src/web_algebra/operations/variable.py index d09274e..9afb42d 100644 --- a/src/web_algebra/operations/variable.py +++ b/src/web_algebra/operations/variable.py @@ -1,5 +1,4 @@ from typing import Any -from mcp import types from web_algebra.operation import Operation @@ -35,8 +34,10 @@ def execute(self, name: str, value: Any, variable_stack: list) -> None: self.set_variable(name, value, variable_stack) return None - def execute_json(self, arguments: dict, variable_stack: list = []) -> None: + def execute_json(self, arguments: dict, variable_stack: list = None) -> None: """JSON execution: evaluate value expression and store variable""" + if variable_stack is None: + variable_stack = [] name: str = arguments["name"] value_expr = arguments["value"] @@ -47,7 +48,3 @@ def execute_json(self, arguments: dict, variable_stack: list = []) -> None: # Call pure function to store the variable return self.execute(name, value, variable_stack) - - def mcp_run(self, arguments: dict, context: Any = None) -> Any: - """MCP execution: plain args → confirmation""" - return [types.TextContent(type="text", text="Variable set successfully")] diff --git a/src/web_algebra/query_source.py b/src/web_algebra/query_source.py new file mode 100644 index 0000000..3dd7137 --- /dev/null +++ b/src/web_algebra/query_source.py @@ -0,0 +1,67 @@ +from typing import Any, Union +from rdflib import Graph, Literal, URIRef +from web_algebra.operation import Operation + + +class QuerySource: + """Mixin for SELECT, CONSTRUCT and DESCRIBE: a query runs over a dataset + given either as `endpoint`, the URI of a SPARQL endpoint, or as `graph`, a + `Graph` in hand (formal-semantics.md §4.3). Exactly one is given. The graph + case is pure and local, so it never touches the client. + """ + + def resolve_source( + self, arguments: dict, variable_stack: list + ) -> Union[URIRef, Graph]: + """Evaluate whichever of `endpoint`/`graph` was given: neither raises + KeyError, both TypeError (§4.3).""" + has_endpoint = "endpoint" in arguments + has_graph = "graph" in arguments + if has_endpoint and has_graph: + raise TypeError( + f"{self.name()} takes exactly one of 'endpoint' and 'graph', got both" + ) + if not has_endpoint and not has_graph: + raise KeyError("endpoint") + + if has_endpoint: + endpoint = Operation.process_json( + self.settings, arguments["endpoint"], self.context, variable_stack + ) + if not isinstance(endpoint, URIRef): + raise TypeError( + f"{self.name()} operation expects 'endpoint' to be URIRef, got {type(endpoint)}" + ) + return endpoint + + data = Operation.process_json( + self.settings, arguments["graph"], self.context, variable_stack + ) + if not isinstance(data, (Graph, dict, list)): + raise TypeError( + f"{self.name()} operation expects 'graph' to be a Graph or RDF data, got {type(data)}" + ) + # an RDF data form is parsed with no base IRI: there is no target to + # resolve against (§2.3, §4.3) + return Operation.to_graph(data) + + def resolve_query(self, arguments: dict, variable_stack: list) -> Literal: + query = Operation.process_json( + self.settings, arguments["query"], self.context, variable_stack + ) + if not Operation.is_string_literal(query): + raise TypeError( + f"{self.name()} operation expects 'query' to be string Literal, got {type(query)}" + ) + return query + + def check_source(self, source: Any, query: Any) -> None: + """Strict type checking before any network I/O (§3.7).""" + if not isinstance(source, (URIRef, Graph)): + raise TypeError( + f"{self.name()} expects its source to be an endpoint URIRef or a Graph, got {type(source).__name__}" + ) + if not Operation.is_string_literal(query): + raise TypeError( + f"{self.name()} expects query to be string Literal, got {type(query).__name__}" + ) diff --git a/src/web_algebra/schema_extraction.py b/src/web_algebra/schema_extraction.py new file mode 100644 index 0000000..4be7e38 --- /dev/null +++ b/src/web_algebra/schema_extraction.py @@ -0,0 +1,124 @@ +import logging +import re +from typing import Any, ClassVar, Optional, Type +from mcp import types +from rdflib import Graph, URIRef +from rdflib.query import Result +from web_algebra.client import SPARQLClient +from web_algebra.client_operation import ClientOperation +from web_algebra.mcp_tool import MCPTool +from web_algebra.operation import Operation +from web_algebra.operations.sparql.values import Values + +# Where an extraction query takes its scope: inside its innermost groups, +# where ?subject is bound. +SCOPE = "%SCOPE%" + +# A named-graph branch of an extraction query, `UNION { GRAPH ?g {`. +_GRAPH_BRANCH = re.compile(r"UNION\s*\{\s*GRAPH\s+\?\w+\s*\{") + + +def without_graph_branches(query: str) -> str: + """The query without its named-graph branches: every + `UNION { GRAPH ?g { ... } }` removed and the default-graph branch beside + it left standing — for an endpoint that refuses the keyword.""" + while (match := _GRAPH_BRANCH.search(query)) is not None: + depth = 0 + i = query.index("{", match.start()) + while i < len(query): + if query[i] == "{": + depth += 1 + elif query[i] == "}": + depth -= 1 + if depth == 0: + break + i += 1 + query = query[: match.start()] + query[i + 1 :] + return query + + +class SchemaExtraction(ClientOperation, MCPTool): + """Mixin for the schema operations (formal-semantics.md §4.6): each takes + `endpoint`, the URI of a SPARQL endpoint, and optionally `bindings`, a + Result whose `subject` column scopes the extraction; it queries the + instance data there and returns an ontology Graph. *Query* effect. + + Not an `Operation` itself, so operation discovery does not register it. + """ + + client_class: ClassVar[Type] = SPARQLClient + + #: The extraction CONSTRUCT query, with `%SCOPE%` where the scope goes. + QUERY: ClassVar[str] = "" + + @classmethod + def inputSchema(cls) -> dict: + return { + "type": "object", + "properties": { + "endpoint": {"type": "string", "description": "SPARQL endpoint URL"}, + "bindings": { + "description": "Optional SELECT result whose ?subject column scopes the extraction to those subjects" + }, + }, + "required": ["endpoint"], + } + + def execute(self, endpoint: URIRef, bindings: Optional[Result] = None) -> Graph: + """Pure function: extract over the endpoint, scoped by `bindings`""" + return self.extract(endpoint, self.scope(endpoint, bindings)) + + def extract(self, endpoint: URIRef, scope: str) -> Graph: + """Run this extraction's query with `scope` (a VALUES block, or "").""" + query = self.QUERY.replace(SCOPE, scope) + if not self.client.takes_graph(str(endpoint)): + query = without_graph_branches(query) + logging.info("%s on %s", self.name(), endpoint) + # POSTed as a form: a scope of a few hundred subjects makes a query that + # a URL cannot carry + return Operation.to_graph(self.client.query(str(endpoint), query, post=True)) + + def scope(self, endpoint: Any, bindings: Optional[Result]) -> str: + """Validate the arguments and render `bindings` as a VALUES block over + ?subject; the empty string when unscoped (§4.6).""" + if not isinstance(endpoint, URIRef): + raise TypeError( + f"{self.name()} operation expects 'endpoint' to be URIRef, got {type(endpoint)}" + ) + if bindings is None: + return "" + if not isinstance(bindings, Result): + raise TypeError( + f"{self.name()} expects 'bindings' to be a Result, got {type(bindings).__name__}" + ) + names = [str(var) for var in (bindings.vars or [])] + if "subject" not in names: + raise ValueError( + f"{self.name()} 'bindings' must bind ?subject, the subjects to describe; it binds {names}" + ) + rows = [ + {str(k): term for k, term in binding.items()} + for binding in (bindings.bindings or []) + ] + # no subjects is a finding, not a schema: an extraction over nothing + # would only pass an empty ontology along + if not rows: + raise ValueError(f"{self.name()}: the 'bindings' matched no subjects") + return Values.render_values(["subject"], rows) + + def execute_json(self, arguments: dict, variable_stack: list = None) -> Graph: + """JSON execution: process arguments with strict type checking""" + endpoint = Operation.process_json( + self.settings, arguments["endpoint"], self.context, variable_stack + ) + bindings = None + if "bindings" in arguments: + bindings = Operation.process_json( + self.settings, arguments["bindings"], self.context, variable_stack + ) + return self.execute(endpoint, bindings) + + def mcp_run(self, arguments: dict, context: Any = None) -> Any: + """MCP execution: plain args → plain results""" + graph = self.execute(URIRef(arguments["endpoint"])) + return [types.TextContent(type="text", text=graph.serialize(format="json-ld"))] diff --git a/tests/SPEC_GAPS.md b/tests/SPEC_GAPS.md index 412579b..cea36f8 100644 --- a/tests/SPEC_GAPS.md +++ b/tests/SPEC_GAPS.md @@ -1,136 +1,143 @@ # Web Algebra spec gaps -This file tracks ambiguities and omissions in `formal-semantics.md` discovered while authoring the test suite. Tests that depend on an unresolved item are marked `pytest.skip("UNCLEAR(spec): ...")` until the spec settles the question. +This file tracks ambiguities and omissions in `formal-semantics.md` discovered while +authoring the test suite. Tests that depend on an unresolved item are marked +`pytest.skip(...)` until the spec settles the question. -Format per entry: -- **Operation / property** — what's unclear - - Assumed for now: ... - - Proposed spec edit: ... +The 2026-07 spec rewrite (evaluation semantics, error taxonomy, per-operation JSON +argument tables) resolved the bulk of the original entries; the resolution record is +kept below so the decisions stay traceable. Only the **Remaining gaps** section is +live. --- -## Operations present in the implementation but absent from the spec catalog - -The spec should either add these to the catalog or the impl should drop them. - -- **Concat** (`operations/string/concat.py`) — no spec entry; signature unknown. - - Assumed for now: no test file authored. - - Proposed spec edit: add to "String Operations" with signature `Sequence Literal → Literal` (or whatever the intended shape is). -- **ExtractOntology** (`operations/schema/extract_ontology.py`) — no spec entry; sibling `Extract*` ops are spec'd. - - Assumed for now: no test file authored. - - Proposed spec edit: add to "Schema Operations" with concrete input/output types. - -(Re-verify the full list during implementation by `ls`-ing `src/web_algebra/operations/` and diffing names against the spec catalog at `formal-semantics.md` lines 57-287.) - -## Result-type and behavior ambiguities - -- **Str** (`Term → Literal`) — what datatype does the result Literal carry? `xsd:string`? simple literal (no datatype)? passthrough for an already-string Literal? SPARQL `STR()` returns a simple literal; spec should pin one. - - Assumed for now: tests assert `isinstance(result, Literal)` and lexical-form equality only; datatype assertions are skipped. - - Proposed spec edit: state result datatype explicitly. -- **URI** (`Term → URI`) — behavior on `Literal` whose lexical form isn't a valid URI? On a `BNode`? Spec lists `BNode` as a `Term` but says nothing about `URI(BNode)`. - - Assumed for now: only URIRef/Literal happy paths exercised; BNode and invalid-URI cases skipped. - - Proposed spec edit: enumerate behavior across all three Term subtypes. -- **EncodeForURI** (`Literal → Literal`) — which character set / RFC? RFC 3986 unreserved? SPARQL `ENCODE_FOR_URI`? They differ on `~`, `*`, etc. - - Assumed for now: only the uncontroversial space-to-`%20` case is asserted. - - Proposed spec edit: cite the RFC or SPARQL function explicitly. -- **Replace** (`Literal × Literal × Literal → Literal`) — regex pattern or literal pattern? SPARQL `REPLACE` is XQuery regex; the existing fixture uses `pattern: "%20"` which works either way. - - Assumed for now: only patterns that are valid as both literal and regex are tested. - - Proposed spec edit: state which. -- **STRUUID** (`() → Literal`) — UUID format? UUID4? Hyphenated? Case? - - Assumed for now: only `isinstance(result, Literal)` and "two consecutive calls differ" are asserted. - - Proposed spec edit: state format. -- **Substitute** (`Literal × Literal × Term → Literal`) — SPARQL variable syntax matched: `?var`, `$var`, or both? How are Term values serialized into the query (URIRef → `<...>`, Literal → `"..."` with datatype? lang tag?). - - Assumed for now: tests skipped pending spec. - - Proposed spec edit: define accepted variable syntax and term serialization rules. -- **Merge** (`Sequence Graph → Graph`) — duplicate triples deduplicated? RDF semantics implies set union; spec is silent. - - Assumed for now: tests assert union behavior; deduplication test marked skip. - - Proposed spec edit: state set vs multiset semantics. -- **Bindings** (`Result → Sequence ResultRow`) — order preservation? Empty Result → empty list? - - Assumed for now: length only is asserted; order assertions skipped. - - Proposed spec edit: state ordering contract. - -## Error semantics - -The Strict Type Checking property (`formal-semantics.md` lines 291-295) says "TypeError raised for mismatched input types" but doesn't extend to other error classes: - -- Missing required argument in JSON dispatch — TypeError? KeyError? ValueError? -- Unknown `@op` — ValueError? Custom exception? -- Live-service operations on network/endpoint failure — propagate? wrap? what type? -- **Variable / Value** lookup on a missing name — error or `None`? - - Assumed for now: tests assert `pytest.raises(Exception)` (broad) for these paths; specific exception class skipped. - - Proposed spec edit: state exception classes. - -## Sequence semantics - -- **ForEach** (`Sequence α × Operation → Sequence β`) — when the inner operation returns `None` or itself a sequence, what's the output shape? Filter `None`s? Flatten? The Sequence Semantics section (lines 302-306) says "Single-item operations applied element-wise" which doesn't answer the multi-item case. - - Assumed for now: only "input length = output length" is asserted, with inner ops that return single items. - - Proposed spec edit: define output shape across each inner-op return shape (None, single, sequence). -- ForEach over a SPARQL `Result` — is iteration order part of the contract? - - Assumed for now: order-sensitive assertions skipped. - -## Filter - -- **Filter signature typo** — `formal-semantics.md` line 99: `(Sequence α × Expression → α) + (Result × Expression → Result)`. The sequence case almost certainly should return `Sequence α`, not `α`. - - Assumed for now: all Filter tests skipped pending correction. - - Proposed spec edit: change `→ α` to `→ Sequence α` in the sequence case. -- **Expression type** — line 27 says `Expression = Operation + Literal + Integer`, but how each kind evaluates as a predicate is undefined. - - Proposed spec edit: define evaluation rules per Expression variant. - -## Variable system - -- **Variable** (`String × Any × VariableStack → ⊥`) — `⊥` (bottom) means non-terminating in type theory; presumably means "no meaningful return". But `execute_json` on the JSON layer must return *something* — what? - - Assumed for now: return value not asserted. - - Proposed spec edit: state JSON-layer return value (`None`? the bound value?). -- **Variable System property** (line 311) is internally contradictory: "Sets variables in current scope, Variable operation manages the stack." Sets-in-current vs manages-the-stack are different operations. - - Proposed spec edit: split into two sentences clarifying which operation is responsible for scope creation vs assignment. -- **Value lookup precedence** — when a name exists both in the variable stack and in the context, which wins? - - Assumed for now: precedence-collision tests skipped. - -## Context system - -- **Current** (`Any → Any`) — behavior when context is unset (default `{}` per the abstract type signature)? Returns the empty dict? Errors? - - Assumed for now: only the "context-set" happy path is tested. -- **Value** — which context container shapes are supported? Spec line 315 says context is `Any` and "varies by operation"; line 318 says Value "accesses context values and variables from stack" without enumerating shapes. The impl supports `ResultRow` (`context[name]`) and any object with `getattr(context, name)`, but not plain `dict`. The default `Operation.context: Any = {}` is a dict, which suggests dict should be valid — but the spec doesn't make that explicit. - - Assumed for now: dict-context test skipped pending spec. - - Proposed spec edit: enumerate the supported context container shapes for Value (ResultRow only? + dict? + arbitrary objects?). -- **Execute** (`Operation → Any`) — narrative description is missing entirely. What does Execute do that JSON dispatch doesn't already? - - Assumed for now: all Execute tests skipped pending spec narrative. - - Proposed spec edit: add narrative description. - -## Schema operations - -- **ExtractClasses / ExtractDatatypeProperties / ExtractObjectProperties** (`URI → Graph`) — what does the URI parameter denote? A SPARQL endpoint, a document URL, or an ontology IRI? Spec narrative is silent. - - Assumed for now: only TypeError-on-non-URIRef case is exercised. - - Proposed spec edit: name the URI's role explicitly. - -## JSON dispatch surface - -The `formal-semantics.md` Execution Architecture section (lines 49-55) declares `execute_json(arguments: dict, variable_stack: list) -> Any` but never specifies the keys that each operation expects in `arguments`. In practice the existing positive fixtures confirm key names for a subset of operations (Str/URI/EncodeForURI: `input`; Replace: `input`/`pattern`/`replacement`; CONSTRUCT: `query`/`endpoint`; PUT: `url`/`data`; ldh-CreateContainer: `parent`/`title`/`slug`; ldh-AddSelect: `url`/`query`/`title`; SPARQLString: `question`). - -The remaining operations (ResolveURI, Merge, Substitute, Variable, Value, Bindings, ForEach, Filter, Execute, GET, POST, PATCH, SELECT, DESCRIBE, schema and most LDH ops) have unverified JSON arg shapes. Tests for those JSON layers are skipped with `UNCLEAR(spec)`. - -- Assumed for now: ForEach uses `{select, operation}` (Python param `select_data` shortened to `select`) — used by `tests/fixtures/positive/for-each-sequence.json`. -- Proposed spec edit: per-operation JSON arg key documentation, or a stated rule (e.g. "JSON arg keys equal Python parameter names"). - -Pending fixtures (will be added once spec confirms key shapes): -- `tests/fixtures/positive/nested-resolve-uri.json` — ResolveURI keys. -- `tests/fixtures/positive/variable-and-value.json` — Variable + Value keys. -- `tests/fixtures/positive/substitute-template.json` — Substitute keys. -- `tests/fixtures/positive/merge-two-graphs.json` — Merge keys. - -## Spec/impl divergences observed on first run - -These four assertions were written from `formal-semantics.md` and failed against the implementation. They are not harness bugs — each is a place where the spec and code disagree, and the team needs to decide which side moves. - -- **`tests/unit/test_str.py::TestStrPure::test_non_term_raises_type_error`** — Spec's Strict Type Checking property mandates TypeError on mismatched input. `Str.execute([1, 2, 3])` returns a Literal instead of raising. Either Str should validate `term` is a `URIRef | Literal | BNode`, or the spec should carve out an exception for Str ("accepts any value, casts via `str(...)`"). -- **`tests/unit/test_select.py::TestSELECTPure::test_wrong_endpoint_type_raises`** — Spec: `URI × Literal → Result`. `SELECT.execute(Literal(...), Literal(...))` does not raise TypeError; it proceeds to an HTTP call. Same conflict between Strict Type Checking and the implementation, scaled to a network side effect. -- **`tests/unit/test_select.py::TestSELECTPure::test_wrong_query_type_raises`** — Same as above with `query=URIRef(...)`. - -## Live-service operations - -- **GET, POST, PUT, PATCH** — return types are spec'd, but content negotiation, headers, status-code handling, redirects, timeouts are all silent. -- **SELECT, CONSTRUCT, DESCRIBE** — same: behavior on 4xx/5xx, malformed query, network failure unspecified. -- **ldh-*** — most return `Any`. What is the meaningful assertion for tests against a live LDH instance? -- **SPARQLString** (`Literal → Literal`) — generates SPARQL "from natural language". Non-deterministic (LLM); no testable invariant beyond return type. - - Assumed for now: pure-layer tests cover input-type validation only; live tests assert return types under the relevant marker. - - Proposed spec edit: define error-handling contract for each I/O op. +## Remaining gaps + +- **`ldh-*` operations** — Appendix A of the spec is informative. It now pins the + return of the *update* operations (the single-row `Result` of §4.4) and makes + them subject to §3.6/§4.4, and `ldh-AddSelect`/`ldh-AddConstruct` are unit-tested + against stubbed HTTP on that basis (return shape, recorded `sp:Select` / + `sp:Construct`, non-2xx → `ValueError`). The remaining `ldh-*` unit tests stay + type-only; their behavior is exercised by the `ldh`-marked integration fixtures + against a live LinkedDataHub. +- **SPARQLString** — §4.3 now pins `endpoint · question · projection · context`. + The type contract and the empty-`context` `ValueError` are tested offline. The + model-dependent contract (parseable simple-literal result, projection honoured, + bounded retry then `ValueError`) needs the model call stubbed, and there is no + infrastructure seam for SPARQLString's model call outside the operation module, + so those tests are skipped. +- **Relative IRIs in a `graph` data form** (§4.3) — the form "is parsed with no + base IRI, so its IRIs must be absolute", but what a relative IRI does (error, + dropped triple, left relative) is unstated. Test skipped (`test_select.py`). +- **Live-service behavior** — §3.7 pins transport failures to + `urllib.error.HTTPError`/`URLError` propagating unwrapped, and §4.3–4.4 pin the + response contract (RDF-only, transparent conneg, non-RDF → `ValueError`); still + unspecified: timeouts, retry policy beyond 429, and redirect handling beyond 308. +- **XPath regex dialect coverage** — `Replace` compiles patterns with Python's `re`. + The common syntax is shared with XPath regular expressions, but XPath-only + constructs (`\p{...}` category escapes, `\i`/`\c`, character-class subtraction + `[a-z-[aeiou]]`) are not supported and surface as `ValueError` (invalid pattern). + This is an implementation gap against the normative fn:replace behavior, not + sanctioned spec behavior. + +## Resolved in the 2026-10 spec revision + +These supersede the corresponding 2026-07 entries below. + +- **Flat sequences** (§3.1, §3.2, §3.8 SEQ): sequences are XDM-flat; Unit and + `Variable` leave no item, a sequence-valued element contributes its items, and a + `Result` is one item (never dissolved). Supersedes "sequence-valued results stay + nested" under *ForEach output shape*. +- **ForEach / Iterate results** (§4.1): the concatenation of the iteration values; + an array `operation` yields the concatenation of its element values per + iteration. Supersedes "operation arrays yield the last non-Unit value". +- **Same-target rule** (§3.6): two iterations of one `ForEach` updating the same + reported URI → `ValueError`, raised when the second write is reported; nested + `ForEach` writes count for every enclosing iteration. +- **Filter by name** (§4.1): a string Literal on a `Binding` looks the variable up + (bare, `?x`, `$x`); a miss → `ValueError`; other pairings → `TypeError`. + Supersedes *Filter signature* below. +- **SELECT/CONSTRUCT/DESCRIBE `graph` operand** (§4.3): exactly one of + `endpoint`/`graph`; neither → `KeyError`, both → `TypeError`; `graph` is pure. +- **Write contract** (§4.4, §3.7): single-row `Result` (`status`, `url` = Location + or effective request URI); non-2xx write → `ValueError`, non-2xx read → + `HTTPError`; `If-Match` from a `HEAD` with the write's `Accept`. +- **Relative `Location`** (§4.4): resolved against the effective request URI + (RFC 3986 §5). +- **SPARQLString empty `context`** (§4.3): the `ValueError` is raised before the + model is called. +- **Schema `bindings`** (§4.6): optional `Result` whose `subject` column scopes the + extraction via `VALUES`; no `subject` / no rows → `ValueError`, non-Result → + `TypeError`. + +## Resolved in the 2026-07 spec revision + +Each item below is now normative in `formal-semantics.md` (section in parentheses), +and the corresponding tests are un-skipped. + +- **Catalog omissions**: `Concat` (§4.2) and `ExtractOntology` (§4.6) added. +- **W3C conformance rule** (§4.2 preamble): operations named after SPARQL 1.1 / + XPath functions follow those definitions *by normative reference*, signatures + included; simple literals are materialized as plain rdflib literals (no + datatype), exactly as rdflib's own SPARQL engine does. +- **Str** (§4.2): per `simple literal STR(literal ltrl)` / `simple literal STR(IRI + rsrc)` — lexical form / codepoint representation as a simple literal; language + tags are not carried over; BNode → `TypeError` (SPARQL type error). +- **Concat** (§4.2): per SPARQL `CONCAT()` result-kind rules — all `xsd:string` → + `xsd:string`; all same language tag → that tag; otherwise simple literal. +- **Replace** (§4.2): per SPARQL `REPLACE()` / `fn:replace` — optional `flags` + argument (`s m i x q`), `$N` capture-group references with `\$`/`\\` escapes, + `err:FORX000*` conditions → `ValueError`, result kind follows the first argument, + and `pattern`/`replacement`/`flags` must be simple literals. +- **URI on BNode / invalid lexical form** (§4.2): BNode → `TypeError`; lexical forms + are not validated against RFC 3986. +- **EncodeForURI character set** (§4.2): percent-encode everything outside the + RFC 3986 unreserved set (`A–Z a–z 0–9 - . _ ~`), per XPath `fn:encode-for-uri`; + result is a simple literal. +- **STRUUID format** (§4.2): simple literal; RFC 4122 version-4, lowercase + hyphenated. +- **Substitute** (§4.3): matches `?var` and `$var` at token boundaries; URI → ``, + Literal → quoted with lang/datatype; BNode → `TypeError`; substitution is textual + and documented as not parse-aware. +- **Merge duplicate semantics** (§4.5): set union; duplicates collapse; blank-node + labels taken as-is (graph union, not RDF merge). +- **Bindings order/empty** (§4.1): order-preserving; empty result → empty sequence. +- **Filter signature** (§4.1): `(Sequence α + Result) × Position → α` — the old + `→ α` was correct for the positional case, which is the only expression kind this + version defines; non-integer expressions → `TypeError`, out-of-range → `ValueError`. +- **ForEach output shape** (§4.1): Unit-valued iterations dropped; sequence-valued + results stay nested; operation arrays yield the last non-Unit value; Result rows + iterate in result order; fresh variable scope per iteration. +- **ForEach pure layer** (§4.1): declared an interpreter-level special form — + `execute_json` only; no pure `execute()` contract. +- **Variable return type** (§4.1): `⊥` corrected to `Unit`; JSON layer returns + `None`; binds in the innermost scope, rebinding overwrites; scope creation belongs + to sequences and ForEach iterations (§3.4), not to Variable. +- **Value context shapes & precedence** (§3.5, §4.1): Binding → bound term, mapping → + member value, other object → attribute; the `$` sigil selects the lookup domain, so + variables and context never shadow; misses → `ValueError`. +- **Current on unset context** (§3.5): `ValueError` — only ForEach establishes a + context. +- **Execute** (§4.1): removed from the algebra — it was the MCP-era entry point + and no document uses it. +- **Extract\* URI role** (§4.6): the URI names a SPARQL endpoint. +- **Error semantics** (§3.7): normative exception table — unknown `@op` → + `ValueError`, type mismatch → `TypeError`, missing required argument → `KeyError`, + unknown variable / context miss → `ValueError`, `null` form → `TypeError`, + transport failures propagate unwrapped. +- **JSON dispatch surface**: every core operation's argument keys are now normative + (§4 catalog, per-entry `JSON:` line); `ldh-*` keys documented informatively + (Appendix A). +- **Strict-typing divergences observed on first run** (Str, SELECT): resolved on the + implementation side — both validate input types before any effect (§3.7). +- **Linked Data response contract** (§4.3–4.4): the HTTP operations are RDF-specific + and symmetric — they read and write RDF graphs; content negotiation is transparent + in the implementation; a non-RDF response (unsupported media type, missing + `Content-Type`, or a body that does not parse as the negotiated format) raises + `ValueError`. +- **Value-domain precision** (§3.1): the former `JSON` summand split into `Object` + (generic-object results, members are Values) and `Data` (RDF data forms, holes are + Terms, the rest raw JSON), with their conversion boundaries stated. +- **Result persistence** (§1.1): Result values are materialized and re-iterable. +- **Value focus-item lookup** (§3.5): closed to `Binding` + mapping; the former + host-reflection (attribute) fallback removed. diff --git a/tests/conftest.py b/tests/conftest.py index 9e3b691..14e6913 100644 --- a/tests/conftest.py +++ b/tests/conftest.py @@ -9,6 +9,7 @@ import json import os +import urllib.request from pathlib import Path from typing import Any @@ -20,6 +21,8 @@ from web_algebra.main import LinkedDataHubSettings, list_operation_subclasses from web_algebra.operation import Operation +from tests.http_stub import StubWeb + @pytest.fixture(scope="session", autouse=True) def _register_operations() -> None: @@ -53,6 +56,25 @@ def fixture_dir() -> Path: return Path(__file__).parent / "fixtures" +@pytest.fixture +def http_stub(monkeypatch) -> StubWeb: + """Answer every urllib request offline and record it. + + Patches ``OpenerDirector.open`` (the urllib boundary every client goes + through), so clients built by operations nested inside other operations + are stubbed too. Set ``http_stub.handler`` to a ``request -> StubResponse`` + callable to script the answers; a non-2xx answer raises ``HTTPError`` as a + real opener does. + """ + web = StubWeb() + + def _open(self, request, data=None, *args, **kwargs): + return web.open(request, data) + + monkeypatch.setattr(urllib.request.OpenerDirector, "open", _open) + return web + + @pytest.fixture def run_op(settings: LinkedDataHubSettings): """Convenience wrapper around Operation.process_json for integration tests.""" diff --git a/tests/http_stub.py b/tests/http_stub.py new file mode 100644 index 0000000..64a8921 --- /dev/null +++ b/tests/http_stub.py @@ -0,0 +1,192 @@ +"""Stub HTTP plumbing for the unit suite (harness, not test cases). + +The HTTP-backed operations talk to the network through ``urllib.request`` +openers. Tests replace ``OpenerDirector.open`` with a ``StubWeb``, so every +client an operation builds (including the clients of operations nested +inside a ``ForEach``) is answered offline, and every request is recorded for +assertions on method, URL, headers and body. + +A real opener raises ``urllib.error.HTTPError`` for a non-2xx answer (its +``HTTPErrorProcessor`` does), so the stub does the same: a handler returning a +non-2xx ``StubResponse`` makes ``open`` raise exactly what urllib would. +""" + +from __future__ import annotations + +import http.client +import io +import json +import re +import urllib.error +import urllib.parse +import urllib.request +from typing import Callable, List, Optional + + +class StubResponse: + """A minimal stand-in for ``http.client.HTTPResponse``.""" + + def __init__( + self, + status: int = 200, + headers: Optional[dict] = None, + body: bytes = b"", + url: Optional[str] = None, + reason: Optional[str] = None, + ): + self.status = status + self.code = status + self.reason = reason or http.client.responses.get(status, "") + self.msg = self.reason + self.headers = http.client.HTTPMessage() + for key, value in (headers or {}).items(): + self.headers[key] = value + self._body = body + self.url = url + + def read(self, *args) -> bytes: + return self._body + + def getcode(self) -> int: + return self.status + + def geturl(self) -> Optional[str]: + return self.url + + def info(self): + return self.headers + + def getheader(self, name, default=None): + return self.headers.get(name, default) + + def getheaders(self): + return list(self.headers.items()) + + def close(self) -> None: + pass + + def __enter__(self): + return self + + def __exit__(self, *exc) -> bool: + return False + + +def request_header(request: urllib.request.Request, name: str) -> Optional[str]: + """Case-insensitive header lookup (urllib capitalizes header names).""" + for key, value in request.header_items(): + if key.lower() == name.lower(): + return value + return None + + +def request_body_text(request: urllib.request.Request) -> str: + data = request.data + if data is None: + return "" + if isinstance(data, bytes): + return data.decode("utf-8", errors="replace") + if isinstance(data, str): + return data + try: + return b"".join(data).decode("utf-8", errors="replace") + except TypeError: + return str(data) + + +def sparql_query_text(request: urllib.request.Request) -> Optional[str]: + """The SPARQL query a request carries — by GET, by POST form, or as a + direct ``application/sparql-query`` POST body — or None.""" + parsed = urllib.parse.urlparse(request.full_url) + params = urllib.parse.parse_qs(parsed.query) + if "query" in params: + return params["query"][0] + if request.get_method() == "POST": + content_type = (request_header(request, "Content-Type") or "").lower() + body = request_body_text(request) + if "application/sparql-query" in content_type: + return body + if "application/x-www-form-urlencoded" in content_type: + form = urllib.parse.parse_qs(body) + if "query" in form: + return form["query"][0] + return None + + +_PROLOGUE = re.compile( + r"^\s*(?:(?:PREFIX\s+[^\s:]*:\s*<[^>]*>|BASE\s+<[^>]*>)\s*)*", re.IGNORECASE +) + + +def query_form(query: str) -> str: + """SELECT / CONSTRUCT / DESCRIBE / ASK — the first keyword after the + prologue (comments are not handled; the harness does not need them).""" + rest = _PROLOGUE.sub("", query, count=1) + match = re.match(r"\s*([A-Za-z]+)", rest) + return match.group(1).upper() if match else "" + + +EMPTY_SELECT = json.dumps({"head": {"vars": []}, "results": {"bindings": []}}).encode() + + +def empty_sparql_answer(request: urllib.request.Request) -> StubResponse: + """An empty, well-formed answer of the query's own form.""" + form = query_form(sparql_query_text(request) or "") + if form == "SELECT": + return StubResponse( + 200, {"Content-Type": "application/sparql-results+json"}, EMPTY_SELECT + ) + if form == "ASK": + return StubResponse( + 200, + {"Content-Type": "application/sparql-results+json"}, + json.dumps({"head": {}, "boolean": False}).encode(), + ) + return StubResponse(200, {"Content-Type": "application/n-triples"}, b"") + + +def default_handler(request: urllib.request.Request) -> StubResponse: + """HEAD: 200, no ETag. SPARQL query: an empty answer of its form. + GET: an empty Turtle graph. Writes: 200, no Location.""" + method = request.get_method() + if sparql_query_text(request) is not None: + return empty_sparql_answer(request) + if method == "HEAD": + return StubResponse(200) + if method == "GET": + return StubResponse(200, {"Content-Type": "text/turtle"}, b"") + return StubResponse(200) + + +class StubWeb: + """Records every request and answers it with ``handler(request)``.""" + + def __init__(self, handler: Optional[Callable] = None): + self.handler: Callable = handler or default_handler + self.requests: List[urllib.request.Request] = [] + + def open(self, request, data=None, *args, **kwargs): + if isinstance(request, str): + request = urllib.request.Request(request, data=data) + self.requests.append(request) + response = self.handler(request) + if response.url is None: + response.url = request.full_url + if not 200 <= response.status < 300: + raise urllib.error.HTTPError( + request.full_url, + response.status, + response.reason, + response.headers, + io.BytesIO(response._body), + ) + return response + + def methods(self) -> List[str]: + return [r.get_method() for r in self.requests] + + def with_method(self, method: str) -> List[urllib.request.Request]: + return [r for r in self.requests if r.get_method() == method] + + def queries(self) -> List[str]: + return [q for q in (sparql_query_text(r) for r in self.requests) if q] diff --git a/tests/unit/test_bindings.py b/tests/unit/test_bindings.py index e502ea5..872af98 100644 --- a/tests/unit/test_bindings.py +++ b/tests/unit/test_bindings.py @@ -1,10 +1,8 @@ -"""Spec: formal-semantics.md "Bindings - Extract binding sequence from SPARQL results" -Abstract: Result → Sequence ResultRow -Python: def execute(self, table: rdflib.query.Result) -> List[Dict[str, Any]] - -Note: the abstract signature says `Sequence ResultRow`, but the Python signature -returns `List[Dict[str, Any]]`. The spec is internally inconsistent here; tests -assert only the abstract sequence shape (length / non-empty / iterable). +"""Spec: formal-semantics.md §4.1 "Bindings — project a SPARQL result to its +row sequence" +Abstract: Result → Sequence Binding +- Order-preserving; an empty result yields the empty sequence. +- A Binding is a partial mapping from variable names to Terms (§1.1). """ from __future__ import annotations @@ -47,12 +45,38 @@ def test_non_result_input_raises(self, settings): with pytest.raises(TypeError): op.execute([1, 2, 3]) - @pytest.mark.skip(reason="UNCLEAR(spec): order preservation not stated") def test_order_preserved(self, settings): - pass + # §4.1: order-preserving + from web_algebra.json_result import JSONResult + + op = Operation.get("Bindings")(settings=settings) + table = JSONResult.from_json( + { + "head": {"vars": ["x"]}, + "results": { + "bindings": [ + {"x": {"type": "literal", "value": v}} + for v in ("a", "b", "c") + ] + }, + } + ) + rows = op.execute(table) + assert [str(row["x"]) for row in rows] == ["a", "b", "c"] class TestBindingsJson: - @pytest.mark.skip(reason="UNCLEAR(spec): JSON arg key for Bindings not given by spec or existing fixtures") def test_json_dispatch(self, settings): - pass + # §4.1 JSON: table: Result + from web_algebra.json_result import JSONResult + + op = Operation.get("Bindings")(settings=settings) + table = JSONResult.from_json( + { + "head": {"vars": ["x"]}, + "results": {"bindings": [{"x": {"type": "literal", "value": "a"}}]}, + } + ) + rows = op.execute_json({"table": table}) + assert len(rows) == 1 + assert rows[0]["x"] == Literal("a") diff --git a/tests/unit/test_client_responses.py b/tests/unit/test_client_responses.py new file mode 100644 index 0000000..ee087e0 --- /dev/null +++ b/tests/unit/test_client_responses.py @@ -0,0 +1,97 @@ +"""Implementation-level tests for the Linked Data response contract +(formal-semantics.md §4.3–4.4): the operations read and write RDF graphs; +content negotiation is transparent; a non-RDF response — unsupported media +type, missing Content-Type, or a body that does not parse as the negotiated +format — raises ValueError. + +These pin the contract at the client seam with a stubbed opener (no network). +""" + +from __future__ import annotations + +import pytest +from rdflib import URIRef + +from web_algebra.client import LinkedDataClient, SPARQLClient + + +class _FakeResponse: + def __init__(self, body: bytes, content_type: str | None): + self._body = body + self.headers = {} if content_type is None else {"Content-Type": content_type} + + def read(self) -> bytes: + return self._body + + +class _FakeOpener: + def __init__(self, response: _FakeResponse): + self._response = response + self.last_request = None + + def open(self, request): + self.last_request = request + return self._response + + +def _get_client(body: bytes, content_type: str | None) -> LinkedDataClient: + client = LinkedDataClient(verify_ssl=False) + client.opener = _FakeOpener(_FakeResponse(body, content_type)) + return client + + +class TestLinkedDataResponses: + def test_rdf_response_parses_to_graph(self): + client = _get_client( + b" .", "text/turtle" + ) + graph = client.get("http://ex/doc") + assert (URIRef("http://ex/s"), URIRef("http://ex/p"), URIRef("http://ex/o")) in graph + + def test_accept_header_offers_rdf_media_types_only(self): + # §4.4: content negotiation is transparent — RDF media types are + # requested + client = _get_client(b"", "text/turtle") + client.get("http://ex/doc") + accept = client.opener.last_request.get_header("Accept") + assert "text/turtle" in accept + assert "application/rdf+xml" in accept + + def test_non_rdf_media_type_raises_value_error(self): + # §4.4: unsupported media type → ValueError + client = _get_client(b"", "text/html") + with pytest.raises(ValueError): + client.get("http://ex/doc") + + def test_missing_content_type_raises_value_error(self): + # §4.4: missing Content-Type → ValueError + client = _get_client(b"anything", None) + with pytest.raises(ValueError): + client.get("http://ex/doc") + + def test_rdf_labelled_garbage_raises_value_error(self): + # §4.4: a body that does not parse as its declared RDF type → + # ValueError + client = _get_client(b"this is not turtle @@@", "text/turtle") + with pytest.raises(ValueError): + client.get("http://ex/doc") + + +class TestSPARQLResponses: + def test_construct_response_not_parsing_as_ntriples_raises(self): + # §4.3: a response that does not parse as the negotiated format → + # ValueError + client = SPARQLClient(verify_ssl=False) + client.opener = _FakeOpener(_FakeResponse(b"not n-triples @@@", None)) + with pytest.raises(ValueError): + client.query( + "http://ex/sparql", + "CONSTRUCT { } WHERE {}", + ) + + def test_select_response_not_parsing_as_json_raises(self): + # §4.3: json.JSONDecodeError is a ValueError subclass + client = SPARQLClient(verify_ssl=False) + client.opener = _FakeOpener(_FakeResponse(b"not json", None)) + with pytest.raises(ValueError): + client.query("http://ex/sparql", "SELECT * WHERE { ?s ?p ?o }") diff --git a/tests/unit/test_concat.py b/tests/unit/test_concat.py new file mode 100644 index 0000000..b98d4ea --- /dev/null +++ b/tests/unit/test_concat.py @@ -0,0 +1,81 @@ +"""Spec: formal-semantics.md §4.2 "Concat — per SPARQL 1.1 CONCAT()": +string literal CONCAT(string literal ltrl1 ... string literal ltrln) +- All inputs typed xsd:string → xsd:string result. +- All inputs carrying the same language tag → result carries it too. +- All other cases (including no inputs) → simple literal. +""" + +from __future__ import annotations + +import pytest +from rdflib import Literal +from rdflib.namespace import XSD + +from web_algebra.operation import Operation + + +class TestConcatPure: + def test_concatenates_lexical_forms(self, settings): + op = Operation.get("Concat")(settings=settings) + result = op.execute([Literal("foo"), Literal("bar")]) + assert str(result) == "foobar" + + def test_all_xsd_string_yields_xsd_string(self, settings): + # §4.2: all inputs typed xsd:string → xsd:string + op = Operation.get("Concat")(settings=settings) + result = op.execute( + [Literal("foo", datatype=XSD.string), Literal("bar", datatype=XSD.string)] + ) + assert result.datatype == XSD.string + + def test_same_language_tag_is_carried(self, settings): + # §4.2: CONCAT("foo"@en, "bar"@en) → "foobar"@en + op = Operation.get("Concat")(settings=settings) + result = op.execute([Literal("foo", lang="en"), Literal("bar", lang="en")]) + assert result.language == "en" + assert str(result) == "foobar" + + def test_mixed_kinds_yield_simple_literal(self, settings): + # §4.2: in all other cases the result is a simple literal + op = Operation.get("Concat")(settings=settings) + result = op.execute([Literal("foo", lang="en"), Literal("bar")]) + assert result.datatype is None and result.language is None + mixed_langs = op.execute([Literal("foo", lang="en"), Literal("bar", lang="de")]) + assert mixed_langs.datatype is None and mixed_langs.language is None + + def test_empty_inputs_yield_empty_simple_literal(self, settings): + op = Operation.get("Concat")(settings=settings) + result = op.execute([]) + assert str(result) == "" + assert result.datatype is None and result.language is None + + def test_non_literal_input_raises_type_error(self, settings): + # §3.7 strict typing + op = Operation.get("Concat")(settings=settings) + with pytest.raises(TypeError): + op.execute([Literal("foo"), 42]) + with pytest.raises(TypeError): + op.execute("not-a-list") + + +class TestConcatJson: + def test_json_dispatch(self, settings): + # §4.2 JSON: inputs: array of string-compatible Literal forms. + # JSON string scalars coerce to xsd:string (§2.2), so the result is + # xsd:string per the all-xsd:string rule. + op = Operation.get("Concat")(settings=settings) + result = op.execute_json({"inputs": ["http://ex/", "x"]}) + assert str(result) == "http://ex/x" + assert result.datatype == XSD.string + + def test_nested_operations_in_inputs(self, settings): + op = Operation.get("Concat")(settings=settings) + result = op.execute_json( + { + "inputs": [ + {"@op": "Str", "args": {"input": {"@id": "http://ex/a"}}}, + "-b", + ] + } + ) + assert str(result) == "http://ex/a-b" diff --git a/tests/unit/test_construct.py b/tests/unit/test_construct.py index 03eace3..e36df47 100644 --- a/tests/unit/test_construct.py +++ b/tests/unit/test_construct.py @@ -1,20 +1,41 @@ -"""Spec: formal-semantics.md "CONSTRUCT - Execute SPARQL CONSTRUCT query" -Abstract: URI × Literal → Graph -Python: def execute(self, endpoint: rdflib.URIRef, query: rdflib.Literal) -> rdflib.Graph +"""Spec: formal-semantics.md §4.3 "CONSTRUCT — execute a SPARQL CONSTRUCT +query over an endpoint or a graph" +Abstract: (URI + Graph) × Literal → Graph +Python: def execute(self, source: URIRef | Graph, query: Literal) -> Graph +JSON: endpoint: URI or graph: Graph (exactly one) · query: string Literal +- Exactly one of endpoint/graph: neither raises KeyError, both TypeError. +- With `graph` the operation is pure and local; with `endpoint` it is a query + effect, the response negotiated as an RDF serialization. +- In JSON, `graph` takes a Graph value or an RDF data form (no base IRI). """ from __future__ import annotations import os +import urllib.error import pytest from rdflib import Graph, Literal, URIRef +from tests.http_stub import StubResponse from web_algebra.operation import Operation +EX = "http://example.org/" + + +def _graph() -> Graph: + g = Graph() + g.add((URIRef(EX + "a"), URIRef(EX + "p"), Literal("1"))) + g.add((URIRef(EX + "b"), URIRef(EX + "p"), Literal("2"))) + return g + + +def _no_network(request): + raise AssertionError(f"unexpected network request: {request.get_method()} {request.full_url}") + class TestCONSTRUCTPure: - def test_wrong_endpoint_type_raises(self, settings): + def test_wrong_source_type_raises(self, settings): op = Operation.get("CONSTRUCT")(settings=settings) with pytest.raises(TypeError): op.execute(Literal("http://example.org/sparql"), Literal("CONSTRUCT { ?s ?p ?o } WHERE { ?s ?p ?o }")) @@ -24,6 +45,47 @@ def test_wrong_query_type_raises(self, settings): with pytest.raises(TypeError): op.execute(URIRef("http://example.org/sparql"), URIRef("CONSTRUCT { ?s ?p ?o } WHERE { ?s ?p ?o }")) + def test_over_graph_returns_constructed_graph(self, settings, http_stub): + # §4.3: with a Graph the query runs locally — pure, no network + http_stub.handler = _no_network + op = Operation.get("CONSTRUCT")(settings=settings) + result = op.execute( + _graph(), + Literal(f"CONSTRUCT {{ ?s <{EX}q> ?o }} WHERE {{ ?s <{EX}p> ?o }}"), + ) + assert isinstance(result, Graph) + assert len(result) == 2 + assert (URIRef(EX + "a"), URIRef(EX + "q"), Literal("1")) in result + assert (URIRef(EX + "b"), URIRef(EX + "q"), Literal("2")) in result + assert http_stub.requests == [] + + def test_wrong_query_type_over_graph_raises(self, settings): + op = Operation.get("CONSTRUCT")(settings=settings) + with pytest.raises(TypeError): + op.execute(_graph(), URIRef("CONSTRUCT WHERE { ?s ?p ?o }")) + + +class TestCONSTRUCTEndpointStubbed: + def test_over_endpoint_returns_graph(self, settings, http_stub): + http_stub.handler = lambda request: StubResponse( + 200, + {"Content-Type": "application/n-triples"}, + f"<{EX}a> <{EX}p> \"1\" .\n".encode(), + ) + op = Operation.get("CONSTRUCT")(settings=settings) + result = op.execute( + URIRef(EX + "sparql"), Literal("CONSTRUCT WHERE { ?s ?p ?o }") + ) + assert isinstance(result, Graph) + assert (URIRef(EX + "a"), URIRef(EX + "p"), Literal("1")) in result + + def test_read_answered_non_2xx_propagates_http_error(self, settings, http_stub): + # §3.7: a read answered outside 2xx → HTTPError, unwrapped + http_stub.handler = lambda request: StubResponse(503) + op = Operation.get("CONSTRUCT")(settings=settings) + with pytest.raises(urllib.error.HTTPError): + op.execute(URIRef(EX + "sparql"), Literal("CONSTRUCT WHERE { ?s ?p ?o }")) + @pytest.mark.sparql class TestCONSTRUCTLive: @@ -40,19 +102,64 @@ def test_returns_graph(self, settings): class TestCONSTRUCTJson: - def test_json_dispatch_arg_shape(self, settings): - # JSON arg keys from existing fixture tests/fixtures/positive/linkeddatahub-put-test.json: - # CONSTRUCT takes {"query": , "endpoint": }. - # We can't dispatch live without an endpoint, but we can validate the call raises - # something other than KeyError when the arg shape is correct. + def test_wrong_endpoint_type_raises_before_network(self, settings, http_stub): + # §3.7: a plain string is a string Literal (§2.2), not a URI + http_stub.handler = _no_network op = Operation.get("CONSTRUCT")(settings=settings) - with pytest.raises(Exception) as exc_info: + with pytest.raises(TypeError): op.execute_json( { - "query": "CONSTRUCT { ?s ?p ?o } WHERE { ?s ?p ?o }", - "endpoint": {"@op": "URI", "args": {"input": "http://127.0.0.1:1/__nope__"}}, + "endpoint": "http://example.org/sparql", + "query": "CONSTRUCT WHERE { ?s ?p ?o }", } ) - # KeyError would mean the arg keys are wrong; any other exception is a plausible - # "endpoint unreachable" path consistent with the spec arg shape. - assert not isinstance(exc_info.value, KeyError) + + def test_neither_endpoint_nor_graph_raises_key_error(self, settings): + # §4.3: neither raises KeyError + op = Operation.get("CONSTRUCT")(settings=settings) + with pytest.raises(KeyError): + op.execute_json({"query": "CONSTRUCT WHERE { ?s ?p ?o }"}) + + def test_both_endpoint_and_graph_raise_type_error(self, settings, http_stub): + # §4.3: both raise TypeError + http_stub.handler = _no_network + op = Operation.get("CONSTRUCT")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + { + "endpoint": {"@id": "http://example.org/sparql"}, + "graph": {"@id": EX + "a", EX + "p": "1"}, + "query": "CONSTRUCT WHERE { ?s ?p ?o }", + } + ) + + def test_graph_as_rdf_data_form(self, settings, http_stub): + # §4.3: in JSON, `graph` takes an RDF data form + http_stub.handler = _no_network + op = Operation.get("CONSTRUCT")(settings=settings) + result = op.execute_json( + { + "graph": {"@id": EX + "a", EX + "p": "1"}, + "query": f"CONSTRUCT {{ ?s <{EX}q> ?o }} WHERE {{ ?s <{EX}p> ?o }}", + } + ) + assert isinstance(result, Graph) + assert set(result) == {(URIRef(EX + "a"), URIRef(EX + "q"), Literal("1"))} + + def test_graph_from_upstream_construct(self, settings, http_stub): + # §4.3: `graph` takes a Graph value — here CONSTRUCT over CONSTRUCT + http_stub.handler = _no_network + op = Operation.get("CONSTRUCT")(settings=settings) + result = op.execute_json( + { + "graph": { + "@op": "CONSTRUCT", + "args": { + "graph": {"@id": EX + "a", EX + "p": "1"}, + "query": f"CONSTRUCT {{ ?s <{EX}q> ?o }} WHERE {{ ?s <{EX}p> ?o }}", + }, + }, + "query": f"CONSTRUCT {{ ?s <{EX}r> ?o }} WHERE {{ ?s <{EX}q> ?o }}", + } + ) + assert set(result) == {(URIRef(EX + "a"), URIRef(EX + "r"), Literal("1"))} diff --git a/tests/unit/test_current.py b/tests/unit/test_current.py index c8ea998..45ba523 100644 --- a/tests/unit/test_current.py +++ b/tests/unit/test_current.py @@ -1,7 +1,7 @@ -"""Spec: formal-semantics.md "Current - Return current context item" -Abstract: Any → Any -Python: def execute(self, current_item: Any) -> Any -Plus Context System property: "Current Operation: Returns the current context item unchanged" (line 317). +"""Spec: formal-semantics.md §4.1 "Current — the context item itself" +Abstract: () → Context +- Yields the context item; raises ValueError when no context is + established (§3.5: only ForEach establishes one). """ from __future__ import annotations @@ -19,9 +19,11 @@ def test_returns_argument_unchanged(self, settings): result = op.execute(sentinel) assert result is sentinel or result == sentinel - @pytest.mark.skip(reason="UNCLEAR(spec): behavior when context is unset (default `{}` per the abstract type signature)") - def test_unset_context(self, settings): - pass + def test_unset_context_raises_value_error(self, settings): + # §3.5/§3.7: no context established → ValueError + op = Operation.get("Current")(settings=settings) + with pytest.raises(ValueError): + op.execute_json({}) class TestCurrentJson: @@ -32,3 +34,13 @@ def test_returns_context_value(self, settings): op = op_cls(settings=settings, context=ctx_value) result = op.execute_json({}) assert result == ctx_value + + def test_yields_the_focus_item(self, settings): + # §3.5: Current yields the focus item itself, not the focus triple + from web_algebra.focus import Focus + + op = Operation.get("Current")( + settings=settings, + context=Focus(item=Literal("the-item"), position=2, size=3), + ) + assert op.execute_json({}) == Literal("the-item") diff --git a/tests/unit/test_describe.py b/tests/unit/test_describe.py index fba2580..a9a9970 100644 --- a/tests/unit/test_describe.py +++ b/tests/unit/test_describe.py @@ -1,6 +1,12 @@ -"""Spec: formal-semantics.md "DESCRIBE - Execute SPARQL DESCRIBE query" -Abstract: URI × Literal → Graph -Python: def execute(self, endpoint: rdflib.URIRef, query: rdflib.Literal) -> rdflib.Graph +"""Spec: formal-semantics.md §4.3 "DESCRIBE — execute a SPARQL DESCRIBE query +over an endpoint or a graph" +Abstract: (URI + Graph) × Literal → Graph +Python: def execute(self, source: URIRef | Graph, query: Literal) -> Graph +JSON: endpoint: URI or graph: Graph (exactly one) · query: string Literal +- Exactly one of endpoint/graph: neither raises KeyError, both TypeError. +- With `graph` the operation is pure and local. +- What a description contains is the query processor's choice (SPARQL 1.1 + §16.4), so only the result type is asserted. """ from __future__ import annotations @@ -12,9 +18,21 @@ from web_algebra.operation import Operation +EX = "http://example.org/" + + +def _graph() -> Graph: + g = Graph() + g.add((URIRef(EX + "a"), URIRef(EX + "p"), Literal("1"))) + return g + + +def _no_network(request): + raise AssertionError(f"unexpected network request: {request.get_method()} {request.full_url}") + class TestDESCRIBEPure: - def test_wrong_endpoint_type_raises(self, settings): + def test_wrong_source_type_raises(self, settings): op = Operation.get("DESCRIBE")(settings=settings) with pytest.raises(TypeError): op.execute(Literal("http://example.org/sparql"), Literal("DESCRIBE ")) @@ -24,6 +42,14 @@ def test_wrong_query_type_raises(self, settings): with pytest.raises(TypeError): op.execute(URIRef("http://example.org/sparql"), URIRef("DESCRIBE ")) + def test_over_graph_returns_graph_without_network(self, settings, http_stub): + # §4.3: with a Graph the query runs locally — pure, no network + http_stub.handler = _no_network + op = Operation.get("DESCRIBE")(settings=settings) + result = op.execute(_graph(), Literal(f"DESCRIBE <{EX}a>")) + assert isinstance(result, Graph) + assert http_stub.requests == [] + @pytest.mark.sparql class TestDESCRIBELive: @@ -37,6 +63,41 @@ def test_returns_graph(self, settings): class TestDESCRIBEJson: - @pytest.mark.skip(reason="UNCLEAR(spec): DESCRIBE JSON arg shape not exemplified by existing fixtures") - def test_json_dispatch(self, settings): - pass + def test_wrong_endpoint_type_raises_before_network(self, settings, http_stub): + # §4.3 JSON: endpoint: URI · query: Literal (xsd:string). + # §3.7: strict typing before any effect. + http_stub.handler = _no_network + op = Operation.get("DESCRIBE")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + { + "endpoint": "http://example.org/sparql", + "query": "DESCRIBE ", + } + ) + + def test_neither_endpoint_nor_graph_raises_key_error(self, settings): + op = Operation.get("DESCRIBE")(settings=settings) + with pytest.raises(KeyError): + op.execute_json({"query": "DESCRIBE "}) + + def test_both_endpoint_and_graph_raise_type_error(self, settings, http_stub): + http_stub.handler = _no_network + op = Operation.get("DESCRIBE")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + { + "endpoint": {"@id": "http://example.org/sparql"}, + "graph": {"@id": EX + "a", EX + "p": "1"}, + "query": f"DESCRIBE <{EX}a>", + } + ) + + def test_graph_as_rdf_data_form(self, settings, http_stub): + http_stub.handler = _no_network + op = Operation.get("DESCRIBE")(settings=settings) + result = op.execute_json( + {"graph": {"@id": EX + "a", EX + "p": "1"}, "query": f"DESCRIBE <{EX}a>"} + ) + assert isinstance(result, Graph) + assert http_stub.requests == [] diff --git a/tests/unit/test_document.py b/tests/unit/test_document.py new file mode 100644 index 0000000..85f3779 --- /dev/null +++ b/tests/unit/test_document.py @@ -0,0 +1,207 @@ +"""Spec: formal-semantics.md §2 (Document Model) and §3 (Evaluation Semantics). + +Covers form discrimination and scalar coercion (§2.2), the URI reference +form, sequence/variable scoping (§3.2, §3.4), and flat sequences (§3.1, +§3.8 SEQ): a sequence's value is the concatenation of its element values — a +sequence-valued element contributes its items, Unit (and so `Variable`) +contributes none, and a Result is one item, never dissolved into rows. +""" + +from __future__ import annotations + +import pytest +from rdflib import Literal, URIRef +from rdflib.namespace import XSD +from rdflib.query import Result + +from web_algebra.json_result import JSONResult +from web_algebra.operation import Operation + + +class TestURIReferenceForm: + def test_id_string_evaluates_to_uri(self, settings): + # §2.2 rule 2: an object whose only member is `@id` evaluates to a URI + result = Operation.process_json(settings, {"@id": "http://example.org/x"}) + assert result == URIRef("http://example.org/x") + + def test_id_with_nested_operation(self, settings): + # §2.2: the inner form may be any form that evaluates to a Term + result = Operation.process_json( + settings, + {"@id": {"@op": "Concat", "args": {"inputs": ["http://ex/", "x"]}}}, + ) + assert result == URIRef("http://ex/x") + + def test_object_with_id_and_other_members_is_rdf_data(self, settings): + # §2.2 rule 3 wins when other members are present: the object is an + # RDF data form and stays a JSON structure + form = {"@id": "http://ex/s", "http://ex/p": "v"} + result = Operation.process_json(settings, form) + assert result == form + + +class TestScalarForms: + def test_string_coerces_to_xsd_string(self, settings): + result = Operation.process_json(settings, "hello") + assert result == Literal("hello", datatype=XSD.string) + + def test_integer_coerces_to_xsd_integer(self, settings): + result = Operation.process_json(settings, 42) + assert result.datatype == XSD.integer + + def test_fractional_number_coerces_to_xsd_double(self, settings): + result = Operation.process_json(settings, 3.14) + assert result.datatype == XSD.double + + def test_boolean_coerces_to_xsd_boolean(self, settings): + result = Operation.process_json(settings, True) + assert result.datatype == XSD.boolean + + def test_null_is_invalid(self, settings): + # §2.2 rule 7 / §3.7: null is not a valid form + with pytest.raises(TypeError): + Operation.process_json(settings, None) + + +class TestOperationCallForm: + def test_unknown_operation_raises_value_error(self, settings): + # §3.7: unknown operation name in `@op` → ValueError + with pytest.raises(ValueError): + Operation.process_json(settings, {"@op": "NoSuchOperation"}) + + def test_execute_is_not_an_operation(self, settings): + # §4 / §5: the former `Execute` operation was removed from the + # algebra, so requesting it is an unknown operation → ValueError + with pytest.raises(ValueError): + Operation.process_json( + settings, + {"@op": "Execute", "args": {"operation": {"@op": "STRUUID"}}}, + ) + + +class TestSequenceScoping: + def test_variable_visible_to_later_steps(self, settings): + # §3.4: a variable bound in a program step is visible to subsequent + # steps of the same sequence + program = [ + {"@op": "Variable", "args": {"name": "x", "value": "v"}}, + {"@op": "Value", "args": {"name": "$x"}}, + ] + result = Operation.process_json(settings, program) + # §3.2: a Variable leaves no item + assert result == [Literal("v", datatype=XSD.string)] + + def test_variable_visible_in_nested_sequence(self, settings): + # §3.4: ...and to forms nested within them + program = [ + {"@op": "Variable", "args": {"name": "x", "value": "v"}}, + [{"@op": "Value", "args": {"name": "$x"}}], + ] + result = Operation.process_json(settings, program) + # §3.1: the nested sequence contributes its items, flat + assert result == [Literal("v", datatype=XSD.string)] + + def test_binding_ceases_after_sequence_ends(self, settings): + # §3.4: the binding ceases to exist after the sequence ends + stack: list = [] + Operation.process_json( + settings, + [{"@op": "Variable", "args": {"name": "x", "value": "v"}}], + variable_stack=stack, + ) + assert stack == [] + with pytest.raises(ValueError): + Operation.process_json( + settings, + {"@op": "Value", "args": {"name": "$x"}}, + variable_stack=stack, + ) + + def test_sequence_value_is_list_of_element_values(self, settings): + # §3.2: the sequence's value is the concatenation of element values + result = Operation.process_json(settings, ["a", 1]) + assert result == [ + Literal("a", datatype=XSD.string), + Literal(1, datatype=XSD.integer), + ] + + +def _table(*values: str) -> JSONResult: + return JSONResult.from_json( + { + "head": {"vars": ["x"]}, + "results": { + "bindings": [{"x": {"type": "literal", "value": v}} for v in values] + }, + } + ) + + +class TestFlatSequences: + def test_nested_sequences_concatenate(self, settings): + # §3.1: a sequence is never an item of a sequence + result = Operation.process_json(settings, [["a", "b"], ["c"]]) + assert result == [ + Literal("a", datatype=XSD.string), + Literal("b", datatype=XSD.string), + Literal("c", datatype=XSD.string), + ] + + def test_deeply_nested_sequences_concatenate(self, settings): + result = Operation.process_json(settings, [[["a"]], [[], "b"]]) + assert result == [ + Literal("a", datatype=XSD.string), + Literal("b", datatype=XSD.string), + ] + + def test_empty_element_sequence_contributes_nothing(self, settings): + result = Operation.process_json(settings, [[], "a", []]) + assert result == [Literal("a", datatype=XSD.string)] + + def test_variable_leaves_no_item(self, settings): + # §3.2: a Variable leaves no item, as xsl:variable leaves none + result = Operation.process_json( + settings, [{"@op": "Variable", "args": {"name": "x", "value": "v"}}] + ) + assert result == [] + + def test_sequence_valued_operation_contributes_its_items(self, settings): + # §3.1: an element whose value is a sequence (a ForEach) contributes + # its items + result = Operation.process_json( + settings, + [ + "a", + { + "@op": "ForEach", + "args": { + "select": ["b", "c"], + "operation": {"@op": "Current", "args": {}}, + }, + }, + ], + ) + assert result == [ + Literal("a", datatype=XSD.string), + Literal("b", datatype=XSD.string), + Literal("c", datatype=XSD.string), + ] + + def test_result_is_not_dissolved(self, settings): + # §3.1: a Result is a value in its own right; concatenation never + # dissolves it into its rows + table = _table("a", "b", "c") + result = Operation.process_json(settings, [table, "z"]) + assert len(result) == 2 + assert isinstance(result[0], Result) + assert len(list(result[0])) == 3 + assert result[1] == Literal("z", datatype=XSD.string) + + def test_bindings_dissolves_explicitly(self, settings): + # §3.1: Bindings does the dissolving, explicitly — its rows become + # items of the enclosing sequence + table = _table("a", "b") + result = Operation.process_json( + settings, [{"@op": "Bindings", "args": {"table": table}}, "z"] + ) + assert len(result) == 3 diff --git a/tests/unit/test_encode_for_uri.py b/tests/unit/test_encode_for_uri.py index 2a2a752..a3175b8 100644 --- a/tests/unit/test_encode_for_uri.py +++ b/tests/unit/test_encode_for_uri.py @@ -31,9 +31,25 @@ def test_uri_input_raises(self, settings): with pytest.raises(TypeError): op.execute(URIRef("http://example.org/x")) - @pytest.mark.skip(reason="UNCLEAR(spec): which character set / RFC — `~`, `*`, `'`, etc. differ across RFC 3986 and SPARQL ENCODE_FOR_URI") - def test_reserved_character_set(self, settings): - pass + def test_rfc3986_unreserved_set_passes_through(self, settings): + # §4.2: everything except A–Z a–z 0–9 - . _ ~ is percent-encoded + op = Operation.get("EncodeForURI")(settings=settings) + result = op.execute(Literal("AZaz09-._~")) + assert str(result) == "AZaz09-._~" + + def test_result_is_simple_literal(self, settings): + # §4.2: simple literal per `simple literal ENCODE_FOR_URI(string literal)` + op = Operation.get("EncodeForURI")(settings=settings) + result = op.execute(Literal("hello world", lang="en")) + assert result.datatype is None and result.language is None + + def test_reserved_characters_are_encoded(self, settings): + # §4.2: reserved characters like / : * ' are encoded (UTF-8) + op = Operation.get("EncodeForURI")(settings=settings) + assert str(op.execute(Literal("a/b"))) == "a%2Fb" + assert str(op.execute(Literal("a:b"))) == "a%3Ab" + assert str(op.execute(Literal("a*b"))) == "a%2Ab" + assert str(op.execute(Literal("a'b"))) == "a%27b" class TestEncodeForURIJson: diff --git a/tests/unit/test_exceptions.py b/tests/unit/test_exceptions.py new file mode 100644 index 0000000..5040dc6 --- /dev/null +++ b/tests/unit/test_exceptions.py @@ -0,0 +1,56 @@ +"""The interpreter's exception taxonomy (src/web_algebra/exceptions.py). + +Each interpreter-level error is a `WebAlgebraError` *and* the built-in the +spec's error table (formal-semantics.md §3.7) mandates — so callers may +`except WebAlgebraError` while the normative built-in contract still holds. +""" + +from __future__ import annotations + +import pytest + +from web_algebra.exceptions import ( + InvalidFormError, + NoFocusError, + UnknownOperationError, + VariableNotFoundError, + WebAlgebraError, +) +from web_algebra.operation import Operation + + +class TestTaxonomyIsBackwardCompatible: + def test_unknown_operation_is_web_algebra_error_and_value_error(self, settings): + # §3.7: unknown `@op` → ValueError + with pytest.raises(UnknownOperationError) as exc: + Operation.process_json(settings, {"@op": "NoSuchOperation"}) + assert isinstance(exc.value, WebAlgebraError) + assert isinstance(exc.value, ValueError) + + def test_null_form_is_web_algebra_error_and_type_error(self, settings): + # §3.7: null form → TypeError + with pytest.raises(InvalidFormError) as exc: + Operation.process_json(settings, None) + assert isinstance(exc.value, WebAlgebraError) + assert isinstance(exc.value, TypeError) + + def test_missing_variable_is_web_algebra_error_and_value_error(self, settings): + # §3.7: unknown variable in `$name` lookup → ValueError + op = Operation.get("Value")(settings=settings) + with pytest.raises(VariableNotFoundError) as exc: + op.execute("$missing", {}, []) + assert isinstance(exc.value, WebAlgebraError) + assert isinstance(exc.value, ValueError) + + def test_no_focus_is_web_algebra_error_and_value_error(self, settings): + # §3.5/§3.7: Current outside an iteration focus → ValueError + op = Operation.get("Current")(settings=settings) + with pytest.raises(NoFocusError) as exc: + op.execute_json({}) + assert isinstance(exc.value, WebAlgebraError) + assert isinstance(exc.value, ValueError) + + def test_web_algebra_error_catches_the_family(self, settings): + # a caller can classify "ill-formed document" with one except clause + with pytest.raises(WebAlgebraError): + Operation.process_json(settings, {"@op": "NoSuchOperation"}) diff --git a/tests/unit/test_execute.py b/tests/unit/test_execute.py deleted file mode 100644 index 8e5b066..0000000 --- a/tests/unit/test_execute.py +++ /dev/null @@ -1,25 +0,0 @@ -"""Spec: formal-semantics.md "Execute - Execute nested operation" -Abstract: Operation → Any -Python: def execute(self, operation: Any) -> Any - -The spec entry has only a signature; no narrative description is given. All -behavioral cases are blocked until the spec adds one. -""" - -from __future__ import annotations - -import pytest - -from web_algebra.operation import Operation - - -class TestExecutePure: - @pytest.mark.skip(reason="UNCLEAR(spec): Execute has no narrative description in formal-semantics.md — what does it do that JSON dispatch doesn't already?") - def test_basic(self, settings): - pass - - -class TestExecuteJson: - @pytest.mark.skip(reason="UNCLEAR(spec): Execute has no narrative description in formal-semantics.md") - def test_json_dispatch(self, settings): - pass diff --git a/tests/unit/test_extract_bindings.py b/tests/unit/test_extract_bindings.py new file mode 100644 index 0000000..50ac16e --- /dev/null +++ b/tests/unit/test_extract_bindings.py @@ -0,0 +1,150 @@ +"""Spec: formal-semantics.md §4.6 "Schema operations" — the optional +`bindings` argument shared by ExtractClasses, ExtractDatatypeProperties, +ExtractObjectProperties and ExtractOntology. +Abstract: URI × Maybe Result → Graph +Python: def execute(self, endpoint: URIRef, + bindings: Optional[Result] = None) -> Graph +JSON: endpoint: URI · bindings: Maybe Result +- Without `bindings` an extraction describes the whole endpoint. +- With `bindings` it describes only the subjects in the result's `subject` + column, put into the extraction query as a VALUES block. +- `bindings` without the variable `subject`, or with no rows → ValueError; + a non-Result → TypeError (before any effect, §3.7). + +The endpoint is stubbed at the urllib boundary (tests/http_stub.py) and +answers every query with an empty result of the query's form. +""" + +from __future__ import annotations + +import pytest +from rdflib import Graph, Literal, URIRef + +from web_algebra.json_result import JSONResult +from web_algebra.operation import Operation + +EX = "http://example.org/" +ENDPOINT = EX + "sparql" +S1, S2 = EX + "s1", EX + "s2" + +EXTRACTIONS = [ + "ExtractClasses", + "ExtractDatatypeProperties", + "ExtractObjectProperties", + "ExtractOntology", +] + + +def _subjects(*iris: str, var: str = "subject") -> JSONResult: + return JSONResult.from_json( + { + "head": {"vars": [var]}, + "results": { + "bindings": [{var: {"type": "uri", "value": iri}} for iri in iris] + }, + } + ) + + +def _assert_values_block(queries: list) -> None: + scoped = [ + q for q in queries if "VALUES" in q.upper() and f"<{S1}>" in q and f"<{S2}>" in q + ] + assert scoped, f"no query carries a VALUES block of the subjects: {queries!r}" + + +@pytest.mark.parametrize("name", EXTRACTIONS) +class TestExtractionWithBindings: + def test_without_bindings_returns_graph(self, name, settings, http_stub): + # §4.6: without bindings, the whole endpoint is described + op = Operation.get(name)(settings=settings) + result = op.execute(URIRef(ENDPOINT)) + assert isinstance(result, Graph) + assert http_stub.queries() + + def test_subjects_appear_as_values_block(self, name, settings, http_stub): + # §4.6: the subjects are put into the extraction query as VALUES + op = Operation.get(name)(settings=settings) + result = op.execute(URIRef(ENDPOINT), _subjects(S1, S2)) + assert isinstance(result, Graph) + _assert_values_block(http_stub.queries()) + + def test_json_subjects_appear_as_values_block(self, name, settings, http_stub): + op = Operation.get(name)(settings=settings) + result = op.execute_json( + {"endpoint": {"@id": ENDPOINT}, "bindings": _subjects(S1, S2)} + ) + assert isinstance(result, Graph) + _assert_values_block(http_stub.queries()) + + def test_other_columns_are_allowed(self, name, settings, http_stub): + # §4.6: only the `subject` column is read + table = JSONResult.from_json( + { + "head": {"vars": ["subject", "label"]}, + "results": { + "bindings": [ + { + "subject": {"type": "uri", "value": iri}, + "label": {"type": "literal", "value": "x"}, + } + for iri in (S1, S2) + ] + }, + } + ) + op = Operation.get(name)(settings=settings) + assert isinstance(op.execute(URIRef(ENDPOINT), table), Graph) + _assert_values_block(http_stub.queries()) + + def test_bindings_without_subject_variable_raise_value_error( + self, name, settings, http_stub + ): + # §4.6/§3.7: bindings that do not bind `subject` → ValueError + op = Operation.get(name)(settings=settings) + with pytest.raises(ValueError): + op.execute(URIRef(ENDPOINT), _subjects(S1, var="s")) + + def test_json_bindings_without_subject_variable_raise_value_error( + self, name, settings, http_stub + ): + op = Operation.get(name)(settings=settings) + with pytest.raises(ValueError): + op.execute_json( + {"endpoint": {"@id": ENDPOINT}, "bindings": _subjects(S1, var="s")} + ) + + def test_bindings_with_no_rows_raise_value_error(self, name, settings, http_stub): + # §4.6/§3.7: bindings with no rows → ValueError + op = Operation.get(name)(settings=settings) + with pytest.raises(ValueError): + op.execute(URIRef(ENDPOINT), _subjects()) + + def test_json_bindings_with_no_rows_raise_value_error( + self, name, settings, http_stub + ): + op = Operation.get(name)(settings=settings) + with pytest.raises(ValueError): + op.execute_json({"endpoint": {"@id": ENDPOINT}, "bindings": _subjects()}) + + @pytest.mark.parametrize( + "bindings", + [Literal("not a result"), URIRef(EX + "x"), [URIRef(S1)]], + ids=["literal", "uri", "list"], + ) + def test_non_result_bindings_raise_type_error( + self, name, bindings, settings, http_stub + ): + # §4.6: a non-Result raises TypeError — before any effect (§3.7) + op = Operation.get(name)(settings=settings) + with pytest.raises(TypeError): + op.execute(URIRef(ENDPOINT), bindings) + assert http_stub.requests == [] + + def test_json_non_result_bindings_raise_type_error( + self, name, settings, http_stub + ): + op = Operation.get(name)(settings=settings) + with pytest.raises(TypeError): + op.execute_json({"endpoint": {"@id": ENDPOINT}, "bindings": "not a result"}) + assert http_stub.requests == [] diff --git a/tests/unit/test_extract_classes.py b/tests/unit/test_extract_classes.py index d71024c..7194446 100644 --- a/tests/unit/test_extract_classes.py +++ b/tests/unit/test_extract_classes.py @@ -1,6 +1,10 @@ """Spec: formal-semantics.md "ExtractClasses - Extract RDF classes from graph" -Abstract: URI → Graph -Python: def execute(self, endpoint: URIRef) -> rdflib.Graph +Abstract: URI × Maybe Result → Graph (§4.6) +Python: def execute(self, endpoint: URIRef, + bindings: Optional[Result] = None) -> rdflib.Graph +- The optional `bindings` contract (subject column → VALUES block, + ValueError/TypeError cases) is covered for all four extractions in + test_extract_bindings.py. """ from __future__ import annotations @@ -17,12 +21,13 @@ def test_wrong_input_type_raises(self, settings): with pytest.raises(TypeError): op.execute(Literal("not-a-uri")) - @pytest.mark.skip(reason="UNCLEAR(spec): is the URI a SPARQL endpoint, document URL, or ontology IRI? — narrative omits this") - def test_happy_path(self, settings): - pass + # §4.6: the URI names a SPARQL endpoint; the happy path queries it and + # is covered under the `sparql` marker via the live suite. class TestExtractClassesJson: - @pytest.mark.skip(reason="UNCLEAR(spec): JSON arg key for ExtractClasses not given by spec or existing fixtures") - def test_json_dispatch(self, settings): - pass + def test_wrong_endpoint_type_raises_before_network(self, settings): + # §4.6 JSON: endpoint: URI. §3.7: strict typing before any effect. + op = Operation.get("ExtractClasses")(settings=settings) + with pytest.raises(TypeError): + op.execute_json({"endpoint": "http://example.org/sparql"}) diff --git a/tests/unit/test_extract_datatype_properties.py b/tests/unit/test_extract_datatype_properties.py index d21dddc..ccbace1 100644 --- a/tests/unit/test_extract_datatype_properties.py +++ b/tests/unit/test_extract_datatype_properties.py @@ -1,6 +1,10 @@ """Spec: formal-semantics.md "ExtractDatatypeProperties - Extract datatype properties from graph" -Abstract: URI → Graph -Python: def execute(self, endpoint: URIRef) -> rdflib.Graph +Abstract: URI × Maybe Result → Graph (§4.6) +Python: def execute(self, endpoint: URIRef, + bindings: Optional[Result] = None) -> rdflib.Graph +- The optional `bindings` contract (subject column → VALUES block, + ValueError/TypeError cases) is covered for all four extractions in + test_extract_bindings.py. """ from __future__ import annotations @@ -17,12 +21,13 @@ def test_wrong_input_type_raises(self, settings): with pytest.raises(TypeError): op.execute(Literal("not-a-uri")) - @pytest.mark.skip(reason="UNCLEAR(spec): is the URI a SPARQL endpoint, document URL, or ontology IRI?") - def test_happy_path(self, settings): - pass + # §4.6: the URI names a SPARQL endpoint; the happy path queries it and + # is covered under the `sparql` marker via the live suite. class TestExtractDatatypePropertiesJson: - @pytest.mark.skip(reason="UNCLEAR(spec): JSON arg key not given by spec or existing fixtures") - def test_json_dispatch(self, settings): - pass + def test_wrong_endpoint_type_raises_before_network(self, settings): + # §4.6 JSON: endpoint: URI. §3.7: strict typing before any effect. + op = Operation.get("ExtractDatatypeProperties")(settings=settings) + with pytest.raises(TypeError): + op.execute_json({"endpoint": "http://example.org/sparql"}) diff --git a/tests/unit/test_extract_object_properties.py b/tests/unit/test_extract_object_properties.py index ef1d8b3..291b73e 100644 --- a/tests/unit/test_extract_object_properties.py +++ b/tests/unit/test_extract_object_properties.py @@ -1,6 +1,10 @@ """Spec: formal-semantics.md "ExtractObjectProperties - Extract object properties from graph" -Abstract: URI → Graph -Python: def execute(self, endpoint: URIRef) -> rdflib.Graph +Abstract: URI × Maybe Result → Graph (§4.6) +Python: def execute(self, endpoint: URIRef, + bindings: Optional[Result] = None) -> rdflib.Graph +- The optional `bindings` contract (subject column → VALUES block, + ValueError/TypeError cases) is covered for all four extractions in + test_extract_bindings.py. """ from __future__ import annotations @@ -17,12 +21,13 @@ def test_wrong_input_type_raises(self, settings): with pytest.raises(TypeError): op.execute(Literal("not-a-uri")) - @pytest.mark.skip(reason="UNCLEAR(spec): is the URI a SPARQL endpoint, document URL, or ontology IRI?") - def test_happy_path(self, settings): - pass + # §4.6: the URI names a SPARQL endpoint; the happy path queries it and + # is covered under the `sparql` marker via the live suite. class TestExtractObjectPropertiesJson: - @pytest.mark.skip(reason="UNCLEAR(spec): JSON arg key not given by spec or existing fixtures") - def test_json_dispatch(self, settings): - pass + def test_wrong_endpoint_type_raises_before_network(self, settings): + # §4.6 JSON: endpoint: URI. §3.7: strict typing before any effect. + op = Operation.get("ExtractObjectProperties")(settings=settings) + with pytest.raises(TypeError): + op.execute_json({"endpoint": "http://example.org/sparql"}) diff --git a/tests/unit/test_extract_ontology.py b/tests/unit/test_extract_ontology.py index 52c9fd7..950d0e4 100644 --- a/tests/unit/test_extract_ontology.py +++ b/tests/unit/test_extract_ontology.py @@ -1,7 +1,11 @@ """Spec: formal-semantics.md "ExtractOntology - Extract a full ontology (classes + datatype + object properties) from a SPARQL endpoint as a single graph" -Abstract: URI → Graph -Python: def execute(self, endpoint: URIRef) -> rdflib.Graph +Abstract: URI × Maybe Result → Graph (§4.6) +Python: def execute(self, endpoint: URIRef, + bindings: Optional[Result] = None) -> rdflib.Graph +- The optional `bindings` contract (subject column → VALUES block, + ValueError/TypeError cases) is covered for all four extractions in + test_extract_bindings.py. """ from __future__ import annotations @@ -18,12 +22,13 @@ def test_wrong_input_type_raises(self, settings): with pytest.raises(TypeError): op.execute(Literal("not-a-uri")) - @pytest.mark.skip(reason="UNCLEAR(spec): is the URI a SPARQL endpoint, document URL, or ontology IRI? — narrative omits this") - def test_happy_path(self, settings): - pass + # §4.6: the URI names a SPARQL endpoint; the happy path queries it and + # is covered under the `sparql` marker via the live suite. class TestExtractOntologyJson: - @pytest.mark.skip(reason="UNCLEAR(spec): JSON arg key for ExtractOntology not given by spec or existing fixtures") - def test_json_dispatch(self, settings): - pass + def test_wrong_endpoint_type_raises_before_network(self, settings): + # §4.6 JSON: endpoint: URI. §3.7: strict typing before any effect. + op = Operation.get("ExtractOntology")(settings=settings) + with pytest.raises(TypeError): + op.execute_json({"endpoint": "http://example.org/sparql"}) diff --git a/tests/unit/test_filter.py b/tests/unit/test_filter.py index 8429391..4b94815 100644 --- a/tests/unit/test_filter.py +++ b/tests/unit/test_filter.py @@ -1,28 +1,194 @@ -"""Spec: formal-semantics.md "Filter - Filter sequences or select from results" -Abstract: (Sequence α × Expression → α) + (Result × Expression → Result) -Python: def execute(self, input_data: Any, expression: Any) -> Union[list, Any] - -The sequence case in the abstract signature has a typo (returns `α`, a single -item, instead of `Sequence α`). Until the spec is corrected and the Expression -semantics are defined (line 27 declares `Expression = Operation + Literal + -Integer` but doesn't define how each kind acts as a predicate), tests are -blocked. +"""Spec: formal-semantics.md §4.1 "Filter — selection by position from a +sequence, or by name from a row, XPath-style" +Abstract: (Sequence α + Result + Binding) × (Position + Literal) → α +- With a Position (xsd:integer Literal): 1-based; a Result input is treated + as its row sequence (yields a Binding). Position < 1 or > length raises + ValueError. +- With a string Literal on a Binding: the term bound to that variable name, + given bare or with `?`/`$`; a miss raises ValueError. +- Any other pairing — a Literal on a sequence or result, a Position on a + Binding, an expression neither integer nor string — raises TypeError. """ from __future__ import annotations import pytest +from rdflib import Literal, URIRef +from rdflib.namespace import XSD +from web_algebra.json_result import JSONResult from web_algebra.operation import Operation -class TestFilterPure: - @pytest.mark.skip(reason="UNCLEAR(spec): line 99 sequence case has typo (`→ α` should be `→ Sequence α`) and Expression evaluation is undefined") - def test_basic(self, settings): - pass +def _result_of(*values: str) -> JSONResult: + return JSONResult.from_json( + { + "head": {"vars": ["x"]}, + "results": { + "bindings": [ + {"x": {"type": "literal", "value": v}} for v in values + ] + }, + } + ) + + +def _pos(n: int) -> Literal: + return Literal(n, datatype=XSD.integer) + + +class TestFilterPositional: + def test_positional_selection_returns_item(self, settings): + # §4.1: 1-based positional selection yields the item itself + op = Operation.get("Filter")(settings=settings) + items = [Literal("a"), Literal("b"), Literal("c")] + assert op.execute(items, _pos(1)) == Literal("a") + assert op.execute(items, _pos(3)) == Literal("c") + + def test_result_input_yields_binding(self, settings): + # §4.1: a Result input is treated as its row sequence + op = Operation.get("Filter")(settings=settings) + row = op.execute(_result_of("a", "b"), _pos(2)) + assert row["x"] == Literal("b") + + def test_position_out_of_range_raises_value_error(self, settings): + # §3.7: Filter position < 1 or > length → ValueError + op = Operation.get("Filter")(settings=settings) + items = [Literal("a")] + with pytest.raises(ValueError): + op.execute(items, _pos(0)) + with pytest.raises(ValueError): + op.execute(items, _pos(2)) + + def test_position_out_of_range_on_result_raises_value_error(self, settings): + op = Operation.get("Filter")(settings=settings) + with pytest.raises(ValueError): + op.execute(_result_of("a"), _pos(2)) + with pytest.raises(ValueError): + op.execute(_result_of("a"), _pos(0)) + + def test_position_on_binding_raises_type_error(self, settings): + # §4.1: a Position on a Binding is not a defined pairing + op = Operation.get("Filter")(settings=settings) + row = op.execute(_result_of("a"), _pos(1)) + with pytest.raises(TypeError): + op.execute(row, _pos(1)) + + +class TestFilterByName: + @pytest.mark.parametrize("name", ["x", "?x", "$x"]) + def test_name_lookup_on_binding(self, settings, name): + # §4.1: a string Literal on a Binding yields the term bound to that + # variable name, bare or with either SPARQL sigil + op = Operation.get("Filter")(settings=settings) + row = op.execute(_result_of("a", "b"), _pos(2)) + assert op.execute(row, Literal(name, datatype=XSD.string)) == Literal("b") + + def test_simple_literal_name_lookup(self, settings): + # §4.2 preamble: a simple literal denotes the same value as xsd:string + op = Operation.get("Filter")(settings=settings) + row = op.execute(_result_of("a"), _pos(1)) + assert op.execute(row, Literal("x")) == Literal("a") + + def test_name_lookup_on_bindings_row(self, settings): + # §1.2: a Binding is a ResultRow or a Dict via Bindings + rows = Operation.get("Bindings")(settings=settings).execute(_result_of("a", "b")) + op = Operation.get("Filter")(settings=settings) + row = op.execute(rows, _pos(1)) + assert op.execute(row, Literal("x")) == Literal("a") + + def test_name_miss_raises_value_error(self, settings): + # §4.1: a miss raises ValueError, as the focus lookup of §3.5 does + op = Operation.get("Filter")(settings=settings) + row = op.execute(_result_of("a"), _pos(1)) + with pytest.raises(ValueError): + op.execute(row, Literal("nope")) + + def test_literal_on_sequence_raises_type_error(self, settings): + op = Operation.get("Filter")(settings=settings) + with pytest.raises(TypeError): + op.execute([Literal("a")], Literal("x")) + + def test_literal_on_result_raises_type_error(self, settings): + op = Operation.get("Filter")(settings=settings) + with pytest.raises(TypeError): + op.execute(_result_of("a"), Literal("x")) + + def test_digit_string_is_a_name_not_a_position(self, settings): + # §2.2: "1" is a string Literal, so on a sequence it is a Literal on a + # sequence → TypeError (only the XML serialization reads it as a + # position, §5) + op = Operation.get("Filter")(settings=settings) + with pytest.raises(TypeError): + op.execute([Literal("a")], Literal("1", datatype=XSD.string)) + + +class TestFilterOtherExpressions: + @pytest.mark.parametrize( + "expression", + [ + URIRef("http://example.org/x"), + Literal(1.0, datatype=XSD.double), + Literal(True, datatype=XSD.boolean), + ], + ids=["uri", "double", "boolean"], + ) + def test_other_expression_raises_type_error(self, settings, expression): + # §4.1: an expression that is neither integer nor string → TypeError + op = Operation.get("Filter")(settings=settings) + with pytest.raises(TypeError): + op.execute([Literal("a")], expression) + row = op.execute(_result_of("a"), _pos(1)) + with pytest.raises(TypeError): + op.execute(row, expression) class TestFilterJson: - @pytest.mark.skip(reason="UNCLEAR(spec): see TestFilterPure") def test_json_dispatch(self, settings): - pass + # §4.1 JSON: a JSON integer coerces to xsd:integer (§2.2), a Position + op = Operation.get("Filter")(settings=settings) + result = op.execute_json({"input": ["a", "b", "c"], "expression": 2}) + # §2.2: the scalar "b" coerces to an xsd:string Literal + assert result == Literal("b", datatype=XSD.string) + + def test_json_position_on_result(self, settings): + op = Operation.get("Filter")(settings=settings) + row = op.execute_json({"input": _result_of("a", "b"), "expression": 1}) + assert row["x"] == Literal("a") + + def test_json_lookup_by_name(self, settings): + # §4.1: Filter(Filter(R, 1), "x") — the first row, then its x + op = Operation.get("Filter")(settings=settings) + result = op.execute_json( + { + "input": { + "@op": "Filter", + "args": {"input": _result_of("a", "b"), "expression": 1}, + }, + "expression": "?x", + } + ) + assert result == Literal("a") + + def test_json_string_on_sequence_raises_type_error(self, settings): + op = Operation.get("Filter")(settings=settings) + with pytest.raises(TypeError): + op.execute_json({"input": ["a"], "expression": "not-an-int"}) + + def test_json_out_of_range_raises_value_error(self, settings): + op = Operation.get("Filter")(settings=settings) + with pytest.raises(ValueError): + op.execute_json({"input": ["a"], "expression": 5}) + + def test_json_name_miss_raises_value_error(self, settings): + op = Operation.get("Filter")(settings=settings) + with pytest.raises(ValueError): + op.execute_json( + { + "input": { + "@op": "Filter", + "args": {"input": _result_of("a"), "expression": 1}, + }, + "expression": "y", + } + ) diff --git a/tests/unit/test_for_each.py b/tests/unit/test_for_each.py index 1b8c80b..8655937 100644 --- a/tests/unit/test_for_each.py +++ b/tests/unit/test_for_each.py @@ -1,45 +1,58 @@ -"""Spec: formal-semantics.md "ForEach - Map operation over sequence (sequence → sequence semantics)" -Abstract: Sequence α × Operation → Sequence β -Python: def execute(self, select_data: Union[List[Any], rdflib.query.Result], - operation: Any) -> List[Any] -Plus Sequence Semantics property (lines 302-306). +"""Spec: formal-semantics.md §4.1 "ForEach — evaluate an operation once per +item of a sequence or per row of a SPARQL result", §3.8 (FOREACH), and the +update rule of §3.6. +Abstract: (Sequence α + Result) × Operation⟨quoted⟩ → Sequence β +- Interpreter-level special form: execute_json only, no pure layer (§4.1). +- Iterates a Sequence item-by-item, a Result row-by-row in result order. +- Each iteration runs in a fresh variable scope under the focus (item i, i, n). +- An array `operation` evaluates as a sequence form within the iteration's + scope: the iteration's value is the concatenation of its element values. +- The result is the concatenation of the iteration values in item order + (§3.1): a sequence-valued iteration contributes its items, a Unit-valued + one nothing; a Result is one item, never dissolved into rows. +- §3.6: two iterations of one ForEach updating the same URI (the URI the + write reports, §4.4) raise ValueError, when the second write is reported; + one iteration may write a URI repeatedly; an update inside a nested ForEach + counts for every enclosing iteration. """ from __future__ import annotations import pytest from rdflib import Literal +from rdflib.namespace import XSD +from rdflib.query import Result +from tests.http_stub import StubResponse, default_handler +from web_algebra.json_result import JSONResult from web_algebra.operation import Operation -class TestForEachPure: - @pytest.mark.skip(reason="UNCLEAR(spec): ForEach pure execute() requires an Operation value plus dispatcher context — abstract signature is testable only via execute_json") - def test_pure(self, settings): - pass +def _str_current() -> dict: + return {"@op": "Str", "args": {"input": {"@op": "Current", "args": {}}}} + + +def _table(*values: str) -> JSONResult: + return JSONResult.from_json( + { + "head": {"vars": ["x"]}, + "results": { + "bindings": [{"x": {"type": "literal", "value": v}} for v in values] + }, + } + ) class TestForEachJson: def test_empty_sequence(self, settings): - # JSON arg keys derived from Python parameter names: select_data → "select" by convention. - # The existing fixture set has no ForEach example; flagged in SPEC_GAPS for confirmation. op = Operation.get("ForEach")(settings=settings) - # Use the JSON dispatcher: an inner Str on each item. - result = op.execute_json( - { - "select": [], - "operation": {"@op": "Str", "args": {"input": {"@op": "Current", "args": {}}}}, - } - ) + result = op.execute_json({"select": [], "operation": _str_current()}) assert result == [] def test_length_matches_input(self, settings): op = Operation.get("ForEach")(settings=settings) result = op.execute_json( - { - "select": ["a", "b", "c"], - "operation": {"@op": "Str", "args": {"input": {"@op": "Current", "args": {}}}}, - } + {"select": ["a", "b", "c"], "operation": _str_current()} ) assert isinstance(result, list) assert len(result) == 3 @@ -47,7 +60,7 @@ def test_length_matches_input(self, settings): assert [str(item) for item in result] == ["a", "b", "c"] def test_non_iterable_select_raises(self, settings): - # Strict Type Checking property: select must be a Sequence or Result. + # §4.1: any other `select` value raises TypeError op = Operation.get("ForEach")(settings=settings) with pytest.raises(TypeError): op.execute_json( @@ -57,10 +70,294 @@ def test_non_iterable_select_raises(self, settings): } ) - @pytest.mark.skip(reason="UNCLEAR(spec): output shape when inner op returns None or a sequence — flatten? filter Nones?") - def test_inner_op_none_handling(self, settings): - pass + def test_unit_valued_iterations_contribute_nothing(self, settings): + # §4.1: a Unit-valued iteration (None) contributes nothing — Variable + # returns Unit, so an all-Variable operation yields the empty sequence. + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + { + "select": ["a", "b"], + "operation": { + "@op": "Variable", + "args": {"name": "x", "value": {"@op": "Current", "args": {}}}, + }, + } + ) + assert result == [] + + def test_sequence_valued_iterations_contribute_their_items(self, settings): + # §4.1/§3.1: the result is the concatenation of the iteration values — + # a nested ForEach's sequence contributes its items, flat. + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + { + "select": ["a", "b"], + "operation": { + "@op": "ForEach", + "args": {"select": ["x", "y"], "operation": _str_current()}, + }, + } + ) + # §4.2: Str returns simple literals + assert result == [Literal("x"), Literal("y"), Literal("x"), Literal("y")] + + def test_operation_array_yields_concatenation_of_its_values(self, settings): + # §4.1: an array operation's value is the concatenation of its element + # values (§3.2) — a Variable leaves no item, every other element does. + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + { + "select": ["a", "b"], + "operation": [ + { + "@op": "Variable", + "args": {"name": "x", "value": {"@op": "Current", "args": {}}}, + }, + {"@op": "Str", "args": {"input": {"@op": "Value", "args": {"name": "$x"}}}}, + "!", + ], + } + ) + bang = Literal("!", datatype=XSD.string) + assert result == [Literal("a"), bang, Literal("b"), bang] + + def test_operation_array_of_units_contributes_nothing(self, settings): + # §3.2: a Variable leaves no item, so an array of Variables is () + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + { + "select": ["a"], + "operation": [ + {"@op": "Variable", "args": {"name": "x", "value": "v"}}, + {"@op": "Variable", "args": {"name": "y", "value": "w"}}, + ], + } + ) + assert result == [] + + def test_result_valued_iterations_are_not_dissolved(self, settings): + # §3.1: a Result is a value in its own right; concatenation never + # dissolves it into its rows + op = Operation.get("ForEach")(settings=settings) + t1, t2 = _table("a", "b"), _table("c") + result = op.execute_json( + {"select": [t1, t2], "operation": {"@op": "Current", "args": {}}} + ) + assert len(result) == 2 + assert all(isinstance(item, Result) for item in result) + assert [len(list(item)) for item in result] == [2, 1] + + def test_iteration_scope_does_not_leak(self, settings): + # §3.4: each iteration runs in a fresh scope — bindings made inside + # do not survive the ForEach. + op = Operation.get("ForEach")(settings=settings) + stack = [{}] + op.execute_json( + { + "select": ["a"], + "operation": [ + {"@op": "Variable", "args": {"name": "x", "value": "v"}}, + {"@op": "Current", "args": {}}, + ], + }, + stack, + ) + assert stack == [{}] + + def test_result_rows_iterate_in_result_order(self, settings): + # §4.1: a Result iterates row-by-row in result order + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + { + "select": _table("a", "b", "c"), + "operation": {"@op": "Value", "args": {"name": "x"}}, + } + ) + assert [str(v) for v in result] == ["a", "b", "c"] + + +# --- §3.6: the xsl:result-document rule for updates inside a ForEach --------- + + +def _put(url: str) -> dict: + return { + "@op": "PUT", + "args": { + "url": {"@id": url}, + "data": {"@id": url, "http://example.org/p": "v"}, + }, + } + + +def _post(url: str) -> dict: + return { + "@op": "POST", + "args": { + "url": {"@id": url}, + "data": {"@id": url, "http://example.org/p": "v"}, + }, + } + + +class TestForEachSameTargetRule: + def test_two_iterations_updating_the_same_uri_raise(self, settings, http_stub): + # §3.6/§3.7: two iterations of one ForEach updating the same URI → + # ValueError + op = Operation.get("ForEach")(settings=settings) + with pytest.raises(ValueError): + op.execute_json( + {"select": ["a", "b"], "operation": _put("http://example.org/doc")} + ) + + def test_second_write_has_been_performed_when_raised(self, settings, http_stub): + # §3.6: the error is raised when the second write is reported, so + # that write has been performed + op = Operation.get("ForEach")(settings=settings) + with pytest.raises(ValueError): + op.execute_json( + {"select": ["a", "b"], "operation": _put("http://example.org/doc")} + ) + assert len(http_stub.with_method("PUT")) == 2 + + def test_iterations_updating_distinct_uris_are_allowed(self, settings, http_stub): + # §3.6: only the same URI twice is an error; each write's Result is + # one item of the concatenation (§3.1) + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + { + "select": ["a", "b"], + "operation": { + "@op": "PUT", + "args": { + "url": { + "@id": { + "@op": "Concat", + "args": { + "inputs": [ + "http://example.org/", + {"@op": "Str", "args": {"input": {"@op": "Current", "args": {}}}}, + ] + }, + } + }, + "data": {"@id": "http://example.org/x", "http://example.org/p": "v"}, + }, + }, + } + ) + assert len(result) == 2 + assert all(isinstance(item, Result) for item in result) + assert len(http_stub.with_method("PUT")) == 2 + + def test_one_iteration_may_write_the_same_uri_twice(self, settings, http_stub): + # §3.6: within one iteration the sequence form orders the writes, so + # a document may be created and then added to + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + { + "select": ["a"], + "operation": [ + _put("http://example.org/doc"), + _post("http://example.org/doc"), + ], + } + ) + assert len(result) == 2 + assert http_stub.methods().count("PUT") == 1 + assert http_stub.methods().count("POST") == 1 + + def test_compared_uri_is_the_reported_one_distinct_locations( + self, settings, http_stub + ): + # §3.6: the URI compared is the one the write reports (§4.4 `url`) — + # POSTs to one container that each create a child (distinct + # Location) do not collide + counter = {"n": 0} + + def handler(request): + if request.get_method() == "POST": + counter["n"] += 1 + return StubResponse( + 201, {"Location": f"http://example.org/container/child{counter['n']}"} + ) + return default_handler(request) + + http_stub.handler = handler + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + {"select": ["a", "b"], "operation": _post("http://example.org/container/")} + ) + assert len(result) == 2 + + def test_compared_uri_is_the_reported_one_same_location(self, settings, http_stub): + # §3.6: writes to different request URIs that report the same URI + # (the response Location, §4.4) collide + def handler(request): + if request.get_method() == "POST": + return StubResponse(201, {"Location": "http://example.org/same"}) + return default_handler(request) + + http_stub.handler = handler + op = Operation.get("ForEach")(settings=settings) + with pytest.raises(ValueError): + op.execute_json( + { + "select": ["a", "b"], + "operation": { + "@op": "POST", + "args": { + "url": { + "@id": { + "@op": "Concat", + "args": { + "inputs": [ + "http://example.org/c/", + {"@op": "Str", "args": {"input": {"@op": "Current", "args": {}}}}, + ] + }, + } + }, + "data": {"@id": "http://example.org/x", "http://example.org/p": "v"}, + }, + }, + } + ) + + def test_nested_for_each_write_counts_for_enclosing_iteration( + self, settings, http_stub + ): + # §3.6: an update made inside a nested ForEach counts for every + # enclosing iteration it runs in — each outer iteration's inner + # ForEach writes the same URI once, so the outer ForEach collides + op = Operation.get("ForEach")(settings=settings) + with pytest.raises(ValueError): + op.execute_json( + { + "select": ["a", "b"], + "operation": { + "@op": "ForEach", + "args": { + "select": ["x"], + "operation": _put("http://example.org/doc"), + }, + }, + } + ) + + def test_separate_for_eaches_may_write_the_same_uri(self, settings, http_stub): + # §3.6: the rule is per ForEach — two ForEach steps of one sequence, + # each writing the URI once, are ordered by the sequence form + program = [ + {"@op": "ForEach", "args": {"select": ["a"], "operation": _put("http://example.org/doc")}}, + {"@op": "ForEach", "args": {"select": ["b"], "operation": _put("http://example.org/doc")}}, + ] + result = Operation.process_json(settings, program) + assert len(result) == 2 + assert len(http_stub.with_method("PUT")) == 2 - @pytest.mark.skip(reason="UNCLEAR(spec): SPARQL Result iteration order") - def test_result_iteration_order(self, settings): - pass + def test_writes_outside_any_for_each_are_unconstrained(self, settings, http_stub): + # §3.6: the rule concerns iterations of a ForEach; a plain sequence + # may write one URI twice + program = [_put("http://example.org/doc"), _put("http://example.org/doc")] + result = Operation.process_json(settings, program) + assert len(result) == 2 diff --git a/tests/unit/test_get.py b/tests/unit/test_get.py index cc3ebd2..3b97998 100644 --- a/tests/unit/test_get.py +++ b/tests/unit/test_get.py @@ -6,10 +6,12 @@ from __future__ import annotations import os +import urllib.error import pytest from rdflib import Graph, Literal, URIRef +from tests.http_stub import StubResponse from web_algebra.operation import Operation @@ -32,6 +34,34 @@ def test_returns_graph(self, settings): class TestGETJson: - @pytest.mark.skip(reason="UNCLEAR(spec): GET JSON arg shape not exemplified by existing fixtures (presumed `{url}`)") - def test_json_dispatch(self, settings): - pass + def test_wrong_url_type_raises_before_network(self, settings): + # §4.4 JSON: url: URI. §3.7: strict typing before any effect — + # a plain string coerces to a string Literal, not a URI. + op = Operation.get("GET")(settings=settings) + with pytest.raises(TypeError): + op.execute_json({"url": "http://example.org/x"}) + + +class TestGETStubbed: + def test_returns_graph(self, settings, http_stub): + http_stub.handler = lambda request: StubResponse( + 200, + {"Content-Type": "text/turtle"}, + b" .", + ) + op = Operation.get("GET")(settings=settings) + result = op.execute(URIRef("http://example.org/doc")) + assert isinstance(result, Graph) + assert len(result) == 1 + assert http_stub.methods() == ["GET"] + + @pytest.mark.parametrize("status", [404, 500]) + def test_read_answered_non_2xx_propagates_http_error( + self, settings, http_stub, status + ): + # §3.7: a read (GET) answered outside 2xx → urllib HTTPError, + # unwrapped — unlike a write, which raises ValueError (§4.4) + http_stub.handler = lambda request: StubResponse(status) + op = Operation.get("GET")(settings=settings) + with pytest.raises(urllib.error.HTTPError): + op.execute(URIRef("http://example.org/doc")) diff --git a/tests/unit/test_iterate.py b/tests/unit/test_iterate.py new file mode 100644 index 0000000..3e9e70d --- /dev/null +++ b/tests/unit/test_iterate.py @@ -0,0 +1,251 @@ +"""Spec: formal-semantics.md §4.1 "Iterate — stateful iteration with +parameter passing between iterations" and §3.8 (ITERATE). +Abstract: (Name ⇀ Value) × Operation⟨quoted⟩ + × Maybe (Name ⇀ Operation⟨quoted⟩) × Maybe Break → Sequence β +- params bound as loop variables; without next-iteration exactly one + iteration runs; next-iteration members are evaluated in the iteration's + environment (loop params + body bindings) and rebind the parameters; + break compares a loop variable's lexical form after rebinding; the + iteration count is capped at 1000 (normative); Iterate establishes no + focus. +- The result is the concatenation of the iteration values, in order (§3.1): + a Unit-valued iteration contributes nothing, a sequence-valued one its + items; an array `operation` yields the concatenation of its element values + per iteration, as in ForEach. +""" + +from __future__ import annotations + +import pytest +from rdflib import Literal + +from web_algebra.operation import Operation + + +def _str_of(name: str) -> dict: + return {"@op": "Str", "args": {"input": {"@op": "Value", "args": {"name": name}}}} + + +class TestIterateJson: + def test_single_iteration_without_next_iteration(self, settings): + # §4.1: without next-iteration, exactly one iteration runs; + # params are bound as variables and read via $name + op = Operation.get("Iterate")(settings=settings) + result = op.execute_json( + {"params": {"url": "https://ex/items"}, "operation": _str_of("$url")} + ) + assert result == [Literal("https://ex/items")] + + def test_next_iteration_rebinds_until_break(self, settings): + # §3.8 (ITERATE): body → next-iteration → rebind → break test + op = Operation.get("Iterate")(settings=settings) + result = op.execute_json( + { + "params": {"s": ""}, + "operation": _str_of("$s"), + "next-iteration": { + "s": { + "@op": "Concat", + "args": { + "inputs": [{"@op": "Value", "args": {"name": "$s"}}, "a"] + }, + } + }, + "break": {"name": "s", "equals": "aaa"}, + } + ) + assert [str(v) for v in result] == ["", "a", "aa"] + + def test_not_equals_break(self, settings): + # §4.1: break with not-equals stops as soon as the variable differs + op = Operation.get("Iterate")(settings=settings) + result = op.execute_json( + { + "params": {"s": "go"}, + "operation": _str_of("$s"), + "next-iteration": {"s": "stop"}, + "break": {"name": "s", "not-equals": "go"}, + } + ) + assert [str(v) for v in result] == ["go"] + + def test_body_bindings_visible_to_next_iteration(self, settings): + # §4.1: next-iteration members are evaluated in the iteration's + # environment — the loop parameters plus the body's bindings + op = Operation.get("Iterate")(settings=settings) + result = op.execute_json( + { + "params": {"url": "a"}, + "operation": [ + { + "@op": "Variable", + "args": { + "name": "page", + "value": { + "@op": "Concat", + "args": { + "inputs": [ + {"@op": "Value", "args": {"name": "$url"}}, + "!", + ] + }, + }, + }, + }, + _str_of("$page"), + ], + "next-iteration": {"url": {"@op": "Value", "args": {"name": "$page"}}}, + "break": {"name": "url", "equals": "a!!!"}, + } + ) + # the Variable element of the array leaves no item (§3.2) + assert [str(v) for v in result] == ["a!", "a!!", "a!!!"] + + def test_array_body_yields_concatenation_per_iteration(self, settings): + # §4.1: array operands as in ForEach — each iteration's value is the + # concatenation of the element values, and the result concatenates + # the iteration values in order + op = Operation.get("Iterate")(settings=settings) + result = op.execute_json( + { + "params": {"s": "a"}, + "operation": [_str_of("$s"), "|"], + "next-iteration": {"s": "b"}, + "break": {"name": "s", "equals": "b"}, + } + ) + assert [str(v) for v in result] == ["a", "|"] + + def test_array_body_two_iterations(self, settings): + # §3.8 (ITERATE): w₁ ⧺ w₂ where each wᵢ is the array's sequence value + op = Operation.get("Iterate")(settings=settings) + result = op.execute_json( + { + "params": {"s": ""}, + "operation": [_str_of("$s"), "|"], + "next-iteration": { + "s": { + "@op": "Concat", + "args": { + "inputs": [{"@op": "Value", "args": {"name": "$s"}}, "a"] + }, + } + }, + "break": {"name": "s", "equals": "aa"}, + } + ) + assert [str(v) for v in result] == ["", "|", "a", "|"] + + def test_sequence_valued_iterations_contribute_their_items(self, settings): + # §3.1: a sequence-valued iteration contributes its items, flat + op = Operation.get("Iterate")(settings=settings) + result = op.execute_json( + { + "params": {"s": "a"}, + "operation": { + "@op": "ForEach", + "args": { + "select": ["x", "y"], + "operation": { + "@op": "Str", + "args": {"input": {"@op": "Current", "args": {}}}, + }, + }, + }, + } + ) + assert result == [Literal("x"), Literal("y")] + + def test_unit_valued_iterations_contribute_nothing(self, settings): + # §4.1/§3.1: a Unit-valued iteration contributes nothing + op = Operation.get("Iterate")(settings=settings) + result = op.execute_json( + { + "params": {"x": "v"}, + "operation": { + "@op": "Variable", + "args": {"name": "y", "value": "w"}, + }, + } + ) + assert result == [] + + def test_loop_scope_does_not_leak(self, settings): + # §3.4: the loop scope (params) and body bindings cease to exist + # after the Iterate + op = Operation.get("Iterate")(settings=settings) + stack: list = [{}] + op.execute_json( + {"params": {"url": "a"}, "operation": _str_of("$url")}, stack + ) + assert stack == [{}] + + def test_iteration_cap_stops_the_loop(self, settings): + # §4.1: the iteration count is bounded by a normative cap of 1000; + # reaching it stops the loop and is not an error + op = Operation.get("Iterate")(settings=settings) + result = op.execute_json( + { + "params": {"s": "x"}, + "operation": _str_of("$s"), + "next-iteration": {"s": {"@op": "Value", "args": {"name": "$s"}}}, + "break": {"name": "s", "equals": "never"}, + } + ) + assert len(result) == 1000 + + def test_enclosing_focus_remains_visible(self, settings): + # §4.1: Iterate establishes no focus; the enclosing one is visible + for_each = Operation.get("ForEach")(settings=settings) + result = for_each.execute_json( + { + "select": ["a"], + "operation": { + "@op": "Iterate", + "args": { + "operation": { + "@op": "Str", + "args": {"input": {"@op": "Current", "args": {}}}, + } + }, + }, + } + ) + # §3.1: the inner Iterate's sequence contributes its items, flat + assert result == [Literal("a")] + + def test_missing_operation_raises_key_error(self, settings): + # §3.7: missing required argument key → KeyError + op = Operation.get("Iterate")(settings=settings) + with pytest.raises(KeyError): + op.execute_json({"params": {"x": "v"}}) + + def test_break_requires_exactly_one_comparison(self, settings): + # §4.1: exactly one of equals/not-equals → ValueError otherwise + op = Operation.get("Iterate")(settings=settings) + with pytest.raises(ValueError): + op.execute_json( + { + "operation": _str_of("$s"), + "params": {"s": "x"}, + "break": {"name": "s"}, + } + ) + with pytest.raises(ValueError): + op.execute_json( + { + "operation": _str_of("$s"), + "params": {"s": "x"}, + "break": {"name": "s", "equals": "a", "not-equals": "b"}, + } + ) + + def test_malformed_arguments_raise_type_error(self, settings): + # §3.7: argument of the wrong type → TypeError + op = Operation.get("Iterate")(settings=settings) + with pytest.raises(TypeError): + op.execute_json({"operation": _str_of("$s"), "params": "not-an-object"}) + with pytest.raises(TypeError): + op.execute_json( + {"operation": _str_of("$s"), "break": "not-an-object"} + ) diff --git a/tests/unit/test_last.py b/tests/unit/test_last.py new file mode 100644 index 0000000..f43e6c9 --- /dev/null +++ b/tests/unit/test_last.py @@ -0,0 +1,36 @@ +"""Spec: formal-semantics.md §4.1 "Last — the size of the iterated sequence, +per XPath fn:last()" +Abstract: () → Literal +- An xsd:integer Literal (§3.5). +- Raises ValueError when no focus is established. +""" + +from __future__ import annotations + +import pytest +from rdflib.namespace import XSD + +from web_algebra.operation import Operation + + +class TestLastJson: + def test_last_is_the_iteration_size_in_every_iteration(self, settings): + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + {"select": ["a", "b", "c"], "operation": {"@op": "Last"}} + ) + assert [int(v) for v in result] == [3, 3, 3] + assert all(v.datatype == XSD.integer for v in result) + + def test_empty_iteration_never_evaluates_last(self, settings): + # ForEach over the empty sequence never evaluates the operand, so + # Last (which would have no meaningful value) is never reached. + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json({"select": [], "operation": {"@op": "Last"}}) + assert result == [] + + def test_no_focus_raises_value_error(self, settings): + # §3.5/§3.7: no focus established → ValueError + op = Operation.get("Last")(settings=settings) + with pytest.raises(ValueError): + op.execute_json({}) diff --git a/tests/unit/test_ldh_add_construct.py b/tests/unit/test_ldh_add_construct.py new file mode 100644 index 0000000..d78337c --- /dev/null +++ b/tests/unit/test_ldh_add_construct.py @@ -0,0 +1,90 @@ +"""Spec: formal-semantics.md Appendix A "ldh-AddConstruct" (informative) +JSON args: url: URI · query: Literal · title: Literal + · description/fragment: Maybe Literal · service: Maybe URI +Python: def execute(self, url: URIRef, query: Literal, title: Literal, description: Literal = None, + fragment: Literal = None, service: URIRef = None) -> Result +- Records a stored query (`sp:Construct`) in the document at `url`. +- An update operation: returns the single-row Result of §4.4 (status, url) + and is subject to the same rules (§3.6, §4.4) — a non-2xx answer to its + write is a ValueError. +""" + +from __future__ import annotations + +import pytest +from rdflib import Literal, URIRef +from rdflib.namespace import XSD +from rdflib.query import Result + +from tests.http_stub import StubResponse, default_handler, request_body_text +from web_algebra.operation import Operation + +DOC = "https://example.org/doc/" +QUERY = "CONSTRUCT WHERE { ?s ?p ?o }" + + +class TestLDHAddConstructPure: + def test_wrong_url_type_raises(self, settings): + op = Operation.get("ldh-AddConstruct")(settings=settings) + with pytest.raises(TypeError): + op.execute( + Literal("not-a-uri"), + Literal(QUERY), + Literal("title"), + ) + + def test_wrong_query_type_raises(self, settings): + op = Operation.get("ldh-AddConstruct")(settings=settings) + with pytest.raises(TypeError): + op.execute( + URIRef("https://example.org/"), + URIRef("not-a-literal"), + Literal("title"), + ) + + +class TestLDHAddConstructStubbed: + def _run(self, settings): + op = Operation.get("ldh-AddConstruct")(settings=settings) + return op.execute_json( + {"url": {"@id": DOC}, "query": QUERY, "title": "Stored query"} + ) + + def test_returns_single_row_write_result(self, settings, http_stub): + # Appendix A: update operations return the single-row Result of §4.4 + result = self._run(settings) + assert isinstance(result, Result) + rows = list(result) + assert len(rows) == 1 + assert rows[0]["status"].datatype == XSD.integer + assert 200 <= int(rows[0]["status"]) < 300 + assert rows[0]["url"] is not None + + def test_records_a_stored_construct(self, settings, http_stub): + # Appendix A: records an sp:Construct in the document + self._run(settings) + bodies = [ + request_body_text(r) + for r in http_stub.requests + if r.get_method() not in ("GET", "HEAD") + ] + assert bodies, "no write was made" + assert any("Construct" in body for body in bodies) + + def test_non_2xx_write_raises_value_error(self, settings, http_stub): + # Appendix A / §4.4 / §3.7: a write answered outside 2xx → ValueError + def handler(request): + if request.get_method() in ("POST", "PUT", "PATCH"): + return StubResponse(403) + return default_handler(request) + + http_stub.handler = handler + with pytest.raises(ValueError): + self._run(settings) + + +@pytest.mark.ldh +class TestLDHAddConstructLive: + @pytest.mark.skip(reason="Live LinkedDataHub run; covered by integration LDH composition fixture instead.") + def test_basic(self, settings_with_auth): + pass diff --git a/tests/unit/test_ldh_add_select.py b/tests/unit/test_ldh_add_select.py index a406d11..0e13b04 100644 --- a/tests/unit/test_ldh_add_select.py +++ b/tests/unit/test_ldh_add_select.py @@ -1,16 +1,27 @@ -"""Spec: formal-semantics.md "ldh-AddSelect - Add SPARQL SELECT service to LinkedDataHub" -Abstract: URI × Literal × Literal × Maybe Literal × Maybe Literal × Maybe URI → Any +"""Spec: formal-semantics.md Appendix A "ldh-AddSelect" (informative) +JSON args: url: URI · query: Literal · title: Literal + · description/fragment: Maybe Literal · service: Maybe URI Python: def execute(self, url: URIRef, query: Literal, title: Literal, description: Literal = None, - fragment: Literal = None, service: URIRef = None) -> Any + fragment: Literal = None, service: URIRef = None) -> Result +- Records a stored query (`sp:Select`) in the document at `url`. +- An update operation: returns the single-row Result of §4.4 (status, url) + and is subject to the same rules (§3.6, §4.4) — a non-2xx answer to its + write is a ValueError. """ from __future__ import annotations import pytest from rdflib import Literal, URIRef +from rdflib.namespace import XSD +from rdflib.query import Result +from tests.http_stub import StubResponse, default_handler, request_body_text from web_algebra.operation import Operation +DOC = "https://example.org/doc/" +QUERY = "SELECT * WHERE { ?s ?p ?o }" + class TestLDHAddSelectPure: def test_wrong_url_type_raises(self, settings): @@ -18,7 +29,7 @@ def test_wrong_url_type_raises(self, settings): with pytest.raises(TypeError): op.execute( Literal("not-a-uri"), - Literal("SELECT * WHERE { ?s ?p ?o }"), + Literal(QUERY), Literal("title"), ) @@ -32,8 +43,48 @@ def test_wrong_query_type_raises(self, settings): ) +class TestLDHAddSelectStubbed: + def _run(self, settings): + op = Operation.get("ldh-AddSelect")(settings=settings) + return op.execute_json( + {"url": {"@id": DOC}, "query": QUERY, "title": "Stored query"} + ) + + def test_returns_single_row_write_result(self, settings, http_stub): + # Appendix A: update operations return the single-row Result of §4.4 + result = self._run(settings) + assert isinstance(result, Result) + rows = list(result) + assert len(rows) == 1 + assert rows[0]["status"].datatype == XSD.integer + assert 200 <= int(rows[0]["status"]) < 300 + assert rows[0]["url"] is not None + + def test_records_a_stored_select(self, settings, http_stub): + # Appendix A: records an sp:Select in the document + self._run(settings) + bodies = [ + request_body_text(r) + for r in http_stub.requests + if r.get_method() not in ("GET", "HEAD") + ] + assert bodies, "no write was made" + assert any("Select" in body for body in bodies) + + def test_non_2xx_write_raises_value_error(self, settings, http_stub): + # Appendix A / §4.4 / §3.7: a write answered outside 2xx → ValueError + def handler(request): + if request.get_method() in ("POST", "PUT", "PATCH"): + return StubResponse(403) + return default_handler(request) + + http_stub.handler = handler + with pytest.raises(ValueError): + self._run(settings) + + @pytest.mark.ldh class TestLDHAddSelectLive: - @pytest.mark.skip(reason="UNCLEAR(spec): return type `Any`. Covered by integration LDH composition fixture instead.") + @pytest.mark.skip(reason="Live LinkedDataHub run; covered by integration LDH composition fixture instead.") def test_basic(self, settings_with_auth): pass diff --git a/tests/unit/test_merge.py b/tests/unit/test_merge.py index cbce6eb..0cdf241 100644 --- a/tests/unit/test_merge.py +++ b/tests/unit/test_merge.py @@ -1,11 +1,11 @@ -"""Spec: formal-semantics.md "Merge - Merge multiple RDF graphs into one" +"""Spec: formal-semantics.md §4.5 "Merge — union of graphs" Abstract: Sequence Graph → Graph -Python: def execute(self, graphs: List[rdflib.Graph]) -> rdflib.Graph +- Set union of triples: duplicate triples collapse. +- JSON: graphs: array of Graph or RDF data forms. """ from __future__ import annotations -import pytest from rdflib import Graph, Literal, URIRef from web_algebra.operation import Operation @@ -44,12 +44,25 @@ def test_two_graphs_union(self, settings): assert t1 in result assert t2 in result - @pytest.mark.skip(reason="UNCLEAR(spec): duplicate-triple semantics (set union vs multiset) not stated") def test_duplicate_triples_deduplicated(self, settings): - pass + # §4.5: set union — duplicate triples collapse + op = Operation.get("Merge")(settings=settings) + triple = (URIRef("http://ex/s"), URIRef("http://ex/p"), Literal("o")) + result = op.execute([_graph_with([triple]), _graph_with([triple])]) + assert len(result) == 1 class TestMergeJson: - @pytest.mark.skip(reason="UNCLEAR(spec): JSON arg key for Merge ('graphs'? 'input'?) not given by spec or existing fixtures") def test_json_dispatch(self, settings): - pass + # §4.5 JSON: graphs: array of Graph or RDF data forms + op = Operation.get("Merge")(settings=settings) + result = op.execute_json( + { + "graphs": [ + {"@id": "http://ex/s1", "http://ex/p": "a"}, + {"@id": "http://ex/s2", "http://ex/p": "b"}, + ] + } + ) + assert (URIRef("http://ex/s1"), URIRef("http://ex/p"), Literal("a")) in result + assert (URIRef("http://ex/s2"), URIRef("http://ex/p"), Literal("b")) in result diff --git a/tests/unit/test_patch.py b/tests/unit/test_patch.py index 6305aca..9c99d90 100644 --- a/tests/unit/test_patch.py +++ b/tests/unit/test_patch.py @@ -1,6 +1,9 @@ """Spec: formal-semantics.md "PATCH - Update RDF data via HTTP PATCH with SPARQL Update" Abstract: URI × Literal → Result Python: def execute(self, url: URIRef, update: Literal) -> Result +- §4.4: returns a single-row Result (status, url); the shared write + contract (url from Location, non-2xx → ValueError, If-Match via HEAD) is + covered for POST, PUT and PATCH in test_write_contract.py. """ from __future__ import annotations @@ -31,6 +34,14 @@ def test_returns_result(self, settings): class TestPATCHJson: - @pytest.mark.skip(reason="UNCLEAR(spec): PATCH JSON arg shape not exemplified by existing fixtures (presumed `{url, update}`)") - def test_json_dispatch(self, settings): - pass + def test_wrong_url_type_raises_before_network(self, settings): + # §4.4 JSON: url: URI · update: Literal (SPARQL Update string). + # §3.7: strict typing before any effect. + op = Operation.get("PATCH")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + { + "url": "http://example.org/x", + "update": "DELETE WHERE { ?s ?p ?o }", + } + ) diff --git a/tests/unit/test_position.py b/tests/unit/test_position.py new file mode 100644 index 0000000..26ebec0 --- /dev/null +++ b/tests/unit/test_position.py @@ -0,0 +1,49 @@ +"""Spec: formal-semantics.md §4.1 "Position — the 1-based position of the +focus item, per XPath fn:position()" +Abstract: () → Literal +- An xsd:integer Literal; within a focus, 1 ≤ position ≤ size (§3.5). +- Raises ValueError when no focus is established. +""" + +from __future__ import annotations + +import pytest +from rdflib.namespace import XSD + +from web_algebra.operation import Operation + + +class TestPositionJson: + def test_positions_run_from_one_to_size(self, settings): + # §3.5: ForEach establishes the focus (item i, i, n) + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + {"select": ["a", "b", "c"], "operation": {"@op": "Position"}} + ) + assert [int(v) for v in result] == [1, 2, 3] + assert all(v.datatype == XSD.integer for v in result) + + def test_nested_for_each_shadows_the_focus(self, settings): + # §3.5: a nested ForEach shadows the outer focus for its operand — + # the inner positions run 1..3 in each of the two outer iterations, + # and the iteration values concatenate flat (§4.1) + op = Operation.get("ForEach")(settings=settings) + result = op.execute_json( + { + "select": ["p", "q"], + "operation": { + "@op": "ForEach", + "args": { + "select": ["x", "y", "z"], + "operation": {"@op": "Position"}, + }, + }, + } + ) + assert [int(v) for v in result] == [1, 2, 3, 1, 2, 3] + + def test_no_focus_raises_value_error(self, settings): + # §3.5/§3.7: no focus established → ValueError + op = Operation.get("Position")(settings=settings) + with pytest.raises(ValueError): + op.execute_json({}) diff --git a/tests/unit/test_post.py b/tests/unit/test_post.py index b75ef9b..59f493b 100644 --- a/tests/unit/test_post.py +++ b/tests/unit/test_post.py @@ -1,6 +1,9 @@ """Spec: formal-semantics.md "POST - Submit RDF data via HTTP POST" Abstract: URI × Graph → Result Python: def execute(self, url: rdflib.URIRef, data: rdflib.Graph) -> Result +- §4.4: returns a single-row Result (status, url); the shared write + contract (url from Location, non-2xx → ValueError, If-Match via HEAD) is + covered for POST, PUT and PATCH in test_write_contract.py. """ from __future__ import annotations @@ -31,6 +34,14 @@ def test_returns_result(self, settings): class TestPOSTJson: - @pytest.mark.skip(reason="UNCLEAR(spec): POST JSON arg shape not exemplified by existing fixtures (presumed `{url, data}`)") - def test_json_dispatch(self, settings): - pass + def test_wrong_url_type_raises_before_network(self, settings): + # §4.4 JSON: url: URI · data: Graph or RDF data form. + # §3.7: strict typing before any effect. + op = Operation.get("POST")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + { + "url": "http://example.org/x", + "data": {"@id": "http://ex/s", "@type": "http://ex/T"}, + } + ) diff --git a/tests/unit/test_process_json_jsonld.py b/tests/unit/test_process_json_jsonld.py index 9fd8d14..e2fafed 100644 --- a/tests/unit/test_process_json_jsonld.py +++ b/tests/unit/test_process_json_jsonld.py @@ -18,7 +18,6 @@ from __future__ import annotations -from types import SimpleNamespace from rdflib import BNode, Graph, Literal, URIRef from rdflib.namespace import RDF @@ -159,15 +158,16 @@ def test_op_in_type_resolves_to_iri(self, settings): assert (DOC_URI, RDF.type, DOC_TYPE) in graph def test_variable_from_context_binding(self, settings): - # Bare name (no $) resolves from the context — the ForEach-row case from - # the original bug report. + # Bare name (no $) resolves from the focus item — the ForEach-row case + # from the original bug report. Item shapes are closed to Binding + + # mapping (formal-semantics.md §3.5), so the row stand-in is a mapping. json_data = { "@id": {"@op": "Value", "args": {"name": "doc"}}, str(FOAF_PRIMARY_TOPIC): { "@id": {"@op": "Value", "args": {"name": "topic"}} }, } - context = SimpleNamespace(doc=DOC_URI, topic=MESSAGE_URI) + context = {"doc": DOC_URI, "topic": MESSAGE_URI} result = Operation.process_json(settings, json_data, context, []) graph = Operation.to_graph(result) diff --git a/tests/unit/test_put.py b/tests/unit/test_put.py index 6d25030..d7b1649 100644 --- a/tests/unit/test_put.py +++ b/tests/unit/test_put.py @@ -1,6 +1,9 @@ """Spec: formal-semantics.md "PUT - Replace RDF data via HTTP PUT" Abstract: URI × Graph → Result Python: def execute(self, url: rdflib.URIRef, data: rdflib.Graph) -> Result +- §4.4: returns a single-row Result (status, url); the shared write + contract (url from Location, non-2xx → ValueError, If-Match via HEAD) is + covered for POST, PUT and PATCH in test_write_contract.py. """ from __future__ import annotations diff --git a/tests/unit/test_replace.py b/tests/unit/test_replace.py index efed367..38a1b75 100644 --- a/tests/unit/test_replace.py +++ b/tests/unit/test_replace.py @@ -1,29 +1,103 @@ -"""Spec: formal-semantics.md "Replace - Replace patterns in strings using regex" -Abstract: Literal × Literal × Literal → Literal -Python: def execute(self, input_str: rdflib.Literal, pattern: rdflib.Literal, - replacement: rdflib.Literal) -> rdflib.Literal -Plus Strict Type Checking property. +"""Spec: formal-semantics.md §4.2 "Replace — per SPARQL 1.1 REPLACE() / +XPath fn:replace": string literal REPLACE(string literal arg, simple literal +pattern, simple literal replacement [, simple literal flags]) +- Replacement string: $N = capture group, \\$ = literal dollar, \\\\ = literal + backslash; other uses of \\ or $ are errors. +- Flags: s, m, i, x, q; invalid flags/pattern/zero-length-matching pattern/ + invalid replacement → ValueError (XPath err:FORX000*). +- pattern/replacement/flags must be simple literals (language tag → TypeError). +- The result is a string literal of the same kind as arg. """ from __future__ import annotations import pytest from rdflib import Literal, URIRef +from rdflib.namespace import XSD from web_algebra.operation import Operation class TestReplacePure: def test_basic_replace(self, settings): - # Pattern that's identical as literal or regex — robust against the regex/literal ambiguity op = Operation.get("Replace")(settings=settings) result = op.execute(Literal("Hello World"), Literal("World"), Literal("Universe")) assert isinstance(result, Literal) assert str(result) == "Hello Universe" + def test_pattern_is_a_regular_expression(self, settings): + # §4.2: XPath fn:replace pattern semantics + op = Operation.get("Replace")(settings=settings) + result = op.execute(Literal("a1b2c3"), Literal("[0-9]"), Literal("#")) + assert str(result) == "a#b#c#" + + def test_group_reference_with_dollar(self, settings): + # §4.2: $N references capture group N — fn:replace example: + # replace("abracadabra", "a(.)", "a$1$1") = "abbraccaddabbra" + op = Operation.get("Replace")(settings=settings) + result = op.execute( + Literal("abracadabra"), Literal("a(.)"), Literal("a$1$1") + ) + assert str(result) == "abbraccaddabbra" + + def test_escaped_dollar_is_literal(self, settings): + # §4.2: \$ is a literal dollar in the replacement + op = Operation.get("Replace")(settings=settings) + result = op.execute(Literal("price"), Literal("price"), Literal("\\$5")) + assert str(result) == "$5" + + def test_bare_dollar_in_replacement_raises(self, settings): + # §4.2: any other use of $ is an error (err:FORX0004) + op = Operation.get("Replace")(settings=settings) + with pytest.raises(ValueError): + op.execute(Literal("abc"), Literal("b"), Literal("x$y")) + + def test_case_insensitive_flag(self, settings): + # §4.2: flags per fn:replace — i + op = Operation.get("Replace")(settings=settings) + result = op.execute(Literal("ABC"), Literal("b"), Literal("x"), Literal("i")) + assert str(result) == "AxC" + + def test_q_flag_treats_pattern_literally(self, settings): + # §4.2: flags per fn:replace — q + op = Operation.get("Replace")(settings=settings) + result = op.execute(Literal("a.b.c"), Literal("."), Literal("x"), Literal("q")) + assert str(result) == "axbxc" + + def test_invalid_flag_raises_value_error(self, settings): + # §3.7/§4.2: invalid flags → ValueError (err:FORX0001) + op = Operation.get("Replace")(settings=settings) + with pytest.raises(ValueError): + op.execute(Literal("abc"), Literal("b"), Literal("x"), Literal("z")) + + def test_zero_length_matching_pattern_raises(self, settings): + # §3.7/§4.2: pattern matching the zero-length string → ValueError + # (err:FORX0003) + op = Operation.get("Replace")(settings=settings) + with pytest.raises(ValueError): + op.execute(Literal("abc"), Literal("b?"), Literal("x")) + + def test_result_kind_follows_first_argument(self, settings): + # §4.2: the result is a string literal of the same kind as arg + op = Operation.get("Replace")(settings=settings) + typed = op.execute( + Literal("chat", datatype=XSD.string), Literal("ch"), Literal("h") + ) + assert typed.datatype == XSD.string + tagged = op.execute(Literal("chat", lang="en"), Literal("ch"), Literal("h")) + assert tagged.language == "en" + assert str(tagged) == "hat" + simple = op.execute(Literal("chat"), Literal("ch"), Literal("h")) + assert simple.datatype is None and simple.language is None + + def test_lang_tagged_pattern_raises_type_error(self, settings): + # §4.2: pattern must be a simple literal per the REPLACE signature + op = Operation.get("Replace")(settings=settings) + with pytest.raises(TypeError): + op.execute(Literal("chat"), Literal("ch", lang="en"), Literal("h")) + def test_uri_input_raises_type_error(self, settings): - # Strict Type Checking property: TypeError on mismatched input. - # Same as existing fixture tests/fixtures/negative/error-case-type-mismatch-uri-to-string.json + # §3.7 strict typing op = Operation.get("Replace")(settings=settings) with pytest.raises(TypeError): op.execute(URIRef("http://example.org/x"), Literal("x"), Literal("y")) @@ -38,14 +112,10 @@ def test_uri_replacement_raises(self, settings): with pytest.raises(TypeError): op.execute(Literal("Hello"), Literal("e"), URIRef("http://example.org/x")) - @pytest.mark.skip(reason="UNCLEAR(spec): regex vs literal pattern semantics — class name and SPARQL parallel suggest regex but spec is silent") - def test_regex_metacharacter_treated_as_regex(self, settings): - pass - class TestReplaceJson: def test_basic_via_json(self, settings): - # JSON arg keys from existing fixture tests/fixtures/positive/complex-operation.json + # §4.2 JSON: input · pattern · replacement · flags (optional) op = Operation.get("Replace")(settings=settings) result = op.execute_json( {"input": "Hello World", "pattern": "World", "replacement": "Universe"} @@ -53,6 +123,13 @@ def test_basic_via_json(self, settings): assert isinstance(result, Literal) assert str(result) == "Hello Universe" + def test_flags_via_json(self, settings): + op = Operation.get("Replace")(settings=settings) + result = op.execute_json( + {"input": "ABC", "pattern": "b", "replacement": "x", "flags": "i"} + ) + assert str(result) == "AxC" + def test_uri_input_raises_via_json(self, settings): op = Operation.get("Replace")(settings=settings) with pytest.raises(TypeError): diff --git a/tests/unit/test_resolve_uri.py b/tests/unit/test_resolve_uri.py index 7d181d4..e1da593 100644 --- a/tests/unit/test_resolve_uri.py +++ b/tests/unit/test_resolve_uri.py @@ -43,12 +43,22 @@ def test_wrong_relative_type_raises(self, settings): with pytest.raises(TypeError): op.execute(URIRef("http://example.org/base/"), URIRef("foo")) - @pytest.mark.skip(reason="UNCLEAR(spec): behavior when relative is itself an absolute URI") - def test_absolute_relative(self, settings): - pass + def test_absolute_relative_returns_itself(self, settings): + # §4.2: RFC 3986 §5 — if `relative` is itself an absolute URI, the + # result is `relative` + op = Operation.get("ResolveURI")(settings=settings) + result = op.execute( + URIRef("http://example.org/base/"), Literal("https://other.example/x") + ) + assert str(result) == "https://other.example/x" class TestResolveURIJson: - @pytest.mark.skip(reason="UNCLEAR(spec): JSON arg key names for ResolveURI not given by spec or existing fixtures") def test_json_dispatch(self, settings): - pass + # §4.2 JSON: base: URI · relative: string-compatible Literal + op = Operation.get("ResolveURI")(settings=settings) + result = op.execute_json( + {"base": {"@id": "http://example.org/base/"}, "relative": "foo"} + ) + assert isinstance(result, URIRef) + assert str(result) == "http://example.org/base/foo" diff --git a/tests/unit/test_select.py b/tests/unit/test_select.py index 4556d18..8afe44d 100644 --- a/tests/unit/test_select.py +++ b/tests/unit/test_select.py @@ -1,29 +1,117 @@ -"""Spec: formal-semantics.md "SELECT - Execute SPARQL SELECT query" -Abstract: URI × Literal → Result -Python: def execute(self, endpoint: rdflib.URIRef, query: rdflib.Literal) -> rdflib.query.Result +"""Spec: formal-semantics.md §4.3 "SELECT — execute a SPARQL SELECT query over +an endpoint or a graph" +Abstract: (URI + Graph) × Literal → Result +Python: def execute(self, source: URIRef | Graph, query: Literal) -> Result +JSON: endpoint: URI or graph: Graph (exactly one) · query: string Literal +- Exactly one of endpoint/graph: neither raises KeyError, both TypeError. +- With `graph` the operation is pure and local; with `endpoint` it is a query + effect. Types are validated before any network I/O (§3.7). +- In JSON, `graph` takes a Graph value or an RDF data form (no base IRI). +- A read answered outside 2xx propagates urllib's HTTPError (§3.7). """ from __future__ import annotations +import json import os +import urllib.error import pytest -from rdflib import Literal, URIRef +from rdflib import Graph, Literal, URIRef +from rdflib.namespace import XSD from rdflib.query import Result +from tests.http_stub import StubResponse from web_algebra.operation import Operation +EX = "http://example.org/" + + +def _graph() -> Graph: + g = Graph() + g.add((URIRef(EX + "a"), URIRef(EX + "p"), Literal("1"))) + g.add((URIRef(EX + "b"), URIRef(EX + "p"), Literal("2"))) + return g + + +def _no_network(request): + raise AssertionError(f"unexpected network request: {request.get_method()} {request.full_url}") + class TestSELECTPure: - def test_wrong_endpoint_type_raises(self, settings): + def test_wrong_source_type_raises(self, settings): op = Operation.get("SELECT")(settings=settings) with pytest.raises(TypeError): - op.execute(Literal("http://example.org/sparql"), Literal("ASK { ?s ?p ?o }")) + op.execute(Literal("http://example.org/sparql"), Literal("SELECT * WHERE { ?s ?p ?o }")) def test_wrong_query_type_raises(self, settings): op = Operation.get("SELECT")(settings=settings) with pytest.raises(TypeError): - op.execute(URIRef("http://example.org/sparql"), URIRef("ASK { ?s ?p ?o }")) + op.execute(URIRef("http://example.org/sparql"), URIRef("SELECT * WHERE { ?s ?p ?o }")) + + def test_wrong_query_type_over_graph_raises(self, settings): + op = Operation.get("SELECT")(settings=settings) + with pytest.raises(TypeError): + op.execute(_graph(), URIRef("SELECT * WHERE { ?s ?p ?o }")) + + def test_over_graph_returns_result(self, settings, http_stub): + # §4.3: with a Graph the query runs locally — pure, no network + http_stub.handler = _no_network + op = Operation.get("SELECT")(settings=settings) + result = op.execute( + _graph(), + Literal(f"SELECT ?s ?o WHERE {{ ?s <{EX}p> ?o }} ORDER BY ?s"), + ) + assert isinstance(result, Result) + rows = list(result) + assert [str(r["s"]) for r in rows] == [EX + "a", EX + "b"] + assert [str(r["o"]) for r in rows] == ["1", "2"] + assert http_stub.requests == [] + + def test_over_graph_accepts_xsd_string_query(self, settings, http_stub): + # §4.3: query is a string Literal, simple or xsd:string + http_stub.handler = _no_network + op = Operation.get("SELECT")(settings=settings) + result = op.execute( + _graph(), + Literal("SELECT ?s WHERE { ?s ?p ?o }", datatype=XSD.string), + ) + assert len(list(result)) == 2 + + def test_over_empty_graph_yields_no_rows(self, settings, http_stub): + http_stub.handler = _no_network + op = Operation.get("SELECT")(settings=settings) + result = op.execute(Graph(), Literal("SELECT ?s WHERE { ?s ?p ?o }")) + assert isinstance(result, Result) + assert len(list(result)) == 0 + + +class TestSELECTEndpointStubbed: + def test_over_endpoint_returns_result(self, settings, http_stub): + # §4.3: with an endpoint, the response is negotiated as SPARQL Results + body = json.dumps( + { + "head": {"vars": ["s"]}, + "results": {"bindings": [{"s": {"type": "uri", "value": EX + "a"}}]}, + } + ).encode() + http_stub.handler = lambda request: StubResponse( + 200, {"Content-Type": "application/sparql-results+json"}, body + ) + op = Operation.get("SELECT")(settings=settings) + result = op.execute(URIRef(EX + "sparql"), Literal("SELECT ?s WHERE { ?s ?p ?o }")) + assert isinstance(result, Result) + assert [r["s"] for r in result] == [URIRef(EX + "a")] + assert len(http_stub.queries()) == 1 + assert "SELECT ?s WHERE { ?s ?p ?o }" in http_stub.queries()[0] + + def test_read_answered_non_2xx_propagates_http_error(self, settings, http_stub): + # §3.7: a read (a SPARQL query) answered outside 2xx → HTTPError, + # unwrapped + http_stub.handler = lambda request: StubResponse(500) + op = Operation.get("SELECT")(settings=settings) + with pytest.raises(urllib.error.HTTPError): + op.execute(URIRef(EX + "sparql"), Literal("SELECT * WHERE { ?s ?p ?o }")) @pytest.mark.sparql @@ -38,6 +126,97 @@ def test_returns_result(self, settings): class TestSELECTJson: - @pytest.mark.skip(reason="UNCLEAR(spec): SELECT JSON arg shape — existing fixtures show `{query, endpoint}` for CONSTRUCT but SELECT is not exemplified") - def test_json_dispatch(self, settings): - pass + def test_wrong_endpoint_type_raises_before_network(self, settings, http_stub): + # §4.3 JSON: endpoint: URI · query: Literal (xsd:string). + # §3.7: TypeError raised before any effect — a plain string is a + # string Literal (§2.2), not a URI. + http_stub.handler = _no_network + op = Operation.get("SELECT")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + { + "endpoint": "http://example.org/sparql", + "query": "SELECT * WHERE { ?s ?p ?o }", + } + ) + + def test_wrong_query_type_raises_before_network(self, settings, http_stub): + http_stub.handler = _no_network + op = Operation.get("SELECT")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + { + "endpoint": {"@id": "http://example.org/sparql"}, + "query": {"@id": "http://example.org/not-a-query"}, + } + ) + + def test_neither_endpoint_nor_graph_raises_key_error(self, settings): + # §4.3: exactly one of endpoint/graph — neither raises KeyError + op = Operation.get("SELECT")(settings=settings) + with pytest.raises(KeyError): + op.execute_json({"query": "SELECT * WHERE { ?s ?p ?o }"}) + + def test_both_endpoint_and_graph_raise_type_error(self, settings, http_stub): + # §4.3: exactly one of endpoint/graph — both raise TypeError + http_stub.handler = _no_network + op = Operation.get("SELECT")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + { + "endpoint": {"@id": "http://example.org/sparql"}, + "graph": {"@id": EX + "a", EX + "p": "1"}, + "query": "SELECT * WHERE { ?s ?p ?o }", + } + ) + + def test_graph_as_rdf_data_form(self, settings, http_stub): + # §4.3: in JSON, `graph` takes an RDF data form, parsed with no base + http_stub.handler = _no_network + op = Operation.get("SELECT")(settings=settings) + result = op.execute_json( + { + "graph": {"@id": EX + "a", EX + "p": "1"}, + "query": f"SELECT ?o WHERE {{ <{EX}a> <{EX}p> ?o }}", + } + ) + assert isinstance(result, Result) + assert [str(r["o"]) for r in result] == ["1"] + assert http_stub.requests == [] + + def test_graph_as_graph_value(self, settings, http_stub): + # §4.3: `graph` takes a Graph value, e.g. the output of Merge + http_stub.handler = _no_network + op = Operation.get("SELECT")(settings=settings) + result = op.execute_json( + { + "graph": { + "@op": "Merge", + "args": { + "graphs": [ + {"@id": EX + "a", EX + "p": "1"}, + {"@id": EX + "b", EX + "p": "2"}, + ] + }, + }, + "query": f"SELECT ?s WHERE {{ ?s <{EX}p> ?o }} ORDER BY ?s", + } + ) + assert [str(r["s"]) for r in result] == [EX + "a", EX + "b"] + + def test_graph_of_wrong_type_raises_type_error(self, settings, http_stub): + # §3.7: graph must be a Graph or RDF data form — a string Literal is + # neither + http_stub.handler = _no_network + op = Operation.get("SELECT")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + {"graph": "not a graph", "query": "SELECT * WHERE { ?s ?p ?o }"} + ) + + def test_relative_iri_in_graph_form(self, settings): + pytest.skip( + "UNCLEAR(spec): §4.3 says a `graph` data form is parsed with no " + "base IRI so its IRIs must be absolute, but not what a relative " + "IRI does (error, dropped triple, or left relative)" + ) diff --git a/tests/unit/test_sparql_string.py b/tests/unit/test_sparql_string.py index f5b3dcf..3b148ae 100644 --- a/tests/unit/test_sparql_string.py +++ b/tests/unit/test_sparql_string.py @@ -1,25 +1,165 @@ -"""Spec: formal-semantics.md "SPARQLString - Generate SPARQL queries from natural language" -Abstract: Literal → Literal -Python: def execute(self, question: Literal) -> Literal +"""Spec: formal-semantics.md §4.3 "SPARQLString — write a SPARQL query for an +endpoint from a natural-language question, via an LLM" +Abstract: URI × Literal × Maybe (Sequence Literal) × Maybe (Sequence Value) + → Literal +Python: def execute(self, endpoint: URIRef, question: Literal, + projection: Optional[List[Literal]] = None, + context: Optional[List[Any]] = None) -> Literal +JSON: endpoint: URI · question: string-compatible Literal + · projection: Maybe (array of string Literals) + · context: Maybe (array of forms) +- Non-deterministic; external service call. Types are checked before any + effect (§3.7), so the type contract is testable offline. +- An empty `context` value (a Result with no rows, an empty Graph) is a + ValueError: the query cannot be written from it. +- The projection / parse retry contract needs the model call stubbed; there + is no infrastructure seam for that, so those cases are skipped. -This operation calls an LLM and is non-deterministic — there is no testable -invariant beyond return type, and even that requires an OpenAI client. +The model is made unreachable (a dummy key and a base URL that refuses +connections), so a test that wrongly reaches the model fails with a +connection error rather than a spec exception. """ from __future__ import annotations import pytest +from rdflib import Graph, Literal, URIRef +from rdflib.namespace import XSD +from web_algebra.json_result import JSONResult from web_algebra.operation import Operation +ENDPOINT = URIRef("http://example.org/sparql") +QUESTION = Literal("Which cities are in Denmark?") + + +@pytest.fixture(autouse=True) +def _unreachable_model(monkeypatch, http_stub): + # harness: keep the operation constructible offline and the model out of + # reach; HTTP to the endpoint is stubbed by http_stub + monkeypatch.setenv("OPENAI_API_KEY", "test-key-not-real") + monkeypatch.setenv("OPENAI_BASE_URL", "http://127.0.0.1:9/v1") + + +def _empty_result() -> JSONResult: + return JSONResult.from_json({"head": {"vars": ["p"]}, "results": {"bindings": []}}) + class TestSPARQLStringPure: - @pytest.mark.skip(reason="UNCLEAR(spec): operation depends on an LLM; no deterministic testable invariant in spec") - def test_basic(self, settings): - pass + def test_endpoint_must_be_uri(self, settings): + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(TypeError): + op.execute(Literal("http://example.org/sparql"), QUESTION) + + @pytest.mark.parametrize( + "question", + [URIRef("http://example.org/q"), Literal(42, datatype=XSD.integer)], + ids=["uri", "integer"], + ) + def test_question_must_be_string_compatible(self, settings, question): + # §4.3: question is a string-compatible Literal (§4.2) + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(TypeError): + op.execute(ENDPOINT, question) + + @pytest.mark.parametrize( + "projection", + [[URIRef("http://example.org/x")], [Literal(1, datatype=XSD.integer)]], + ids=["uri-item", "integer-item"], + ) + def test_projection_items_must_be_string_literals(self, settings, projection): + # §4.3: projection is an array of string Literals (variable names) + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(TypeError): + op.execute(ENDPOINT, QUESTION, projection) + + def test_empty_result_context_raises_value_error(self, settings): + # §4.3/§3.7: a context value that is a Result with no rows → ValueError + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(ValueError): + op.execute(ENDPOINT, QUESTION, None, [_empty_result()]) + + def test_empty_graph_context_raises_value_error(self, settings): + # §4.3/§3.7: a context value that is an empty Graph → ValueError + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(ValueError): + op.execute(ENDPOINT, QUESTION, None, [Graph()]) class TestSPARQLStringJson: - @pytest.mark.skip(reason="UNCLEAR(spec): same as TestSPARQLStringPure") - def test_json_dispatch(self, settings): + def test_missing_endpoint_raises_key_error(self, settings): + # §3.7: missing required argument → KeyError + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(KeyError): + op.execute_json({"question": "Which cities are in Denmark?"}) + + def test_missing_question_raises_key_error(self, settings): + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(KeyError): + op.execute_json({"endpoint": {"@id": str(ENDPOINT)}}) + + def test_plain_string_endpoint_raises_type_error(self, settings): + # §2.2: a plain string is a string Literal, never a URI + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + {"endpoint": "http://example.org/sparql", "question": "Which cities?"} + ) + + def test_uri_question_raises_type_error(self, settings): + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + { + "endpoint": {"@id": str(ENDPOINT)}, + "question": {"@id": "http://example.org/q"}, + } + ) + + def test_non_string_projection_item_raises_type_error(self, settings): + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(TypeError): + op.execute_json( + { + "endpoint": {"@id": str(ENDPOINT)}, + "question": "Which cities?", + "projection": ["city", 7], + } + ) + + def test_empty_context_value_raises_value_error(self, settings): + # §4.3: context forms are evaluated eagerly; an empty Result → ValueError + op = Operation.get("SPARQLString")(settings=settings) + with pytest.raises(ValueError): + op.execute_json( + { + "endpoint": {"@id": str(ENDPOINT)}, + "question": "Which cities?", + "context": [_empty_result()], + } + ) + + +class TestSPARQLStringModelContract: + @pytest.mark.skip( + reason="§4.3: result is a simple literal holding a parseable SPARQL 1.1 " + "query; needs the model call stubbed, and there is no infrastructure " + "seam for SPARQLString's model call outside the operation module" + ) + def test_result_is_parseable_simple_literal(self, settings): + pass + + @pytest.mark.skip( + reason="§4.3: with projection the query is a SELECT projecting every " + "named variable; needs a stubbed model (no infrastructure seam)" + ) + def test_projection_is_honoured(self, settings): + pass + + @pytest.mark.skip( + reason="§4.3: an unparseable / wrongly-projected answer goes back to " + "the model a bounded number of times, then ValueError; needs a " + "stubbed model (no infrastructure seam)" + ) + def test_retry_then_value_error(self, settings): pass diff --git a/tests/unit/test_str.py b/tests/unit/test_str.py index 96c157b..0875a04 100644 --- a/tests/unit/test_str.py +++ b/tests/unit/test_str.py @@ -1,7 +1,9 @@ -"""Spec: formal-semantics.md "Str - Convert any term to string literal" -Abstract: Term → Literal -Python: def execute(self, term: rdflib.term.Node) -> rdflib.Literal -Plus the Strict Type Checking property (lines 291-295). +"""Spec: formal-semantics.md §4.2 "Str — the lexical form of a Term, per +SPARQL 1.1 STR()": simple literal STR(literal ltrl) / simple literal STR(IRI rsrc) +- Returns the lexical form / codepoint representation as a simple literal + (no datatype, no language tag). +- The language tag is not carried over. +- BNode raises TypeError (a SPARQL type error), as does any non-Term. """ from __future__ import annotations @@ -12,43 +14,61 @@ from web_algebra.operation import Operation +def _is_simple_literal(value) -> bool: + return ( + isinstance(value, Literal) + and value.datatype is None + and value.language is None + ) + + class TestStrPure: - def test_uri_returns_literal(self, settings): + def test_uri_yields_simple_literal_of_codepoints(self, settings): + # §4.2: simple literal STR(IRI rsrc) op = Operation.get("Str")(settings=settings) result = op.execute(URIRef("http://example.org/foo")) - assert isinstance(result, Literal) + assert _is_simple_literal(result) assert str(result) == "http://example.org/foo" - def test_literal_returns_literal(self, settings): + def test_plain_literal_yields_simple_literal(self, settings): op = Operation.get("Str")(settings=settings) result = op.execute(Literal("hello")) - assert isinstance(result, Literal) + assert _is_simple_literal(result) assert str(result) == "hello" - def test_bnode_returns_literal(self, settings): + def test_lang_tag_is_not_carried_over(self, settings): + # §4.2: as in SPARQL, STR("hallo"@de) → "hallo" (simple literal) + op = Operation.get("Str")(settings=settings) + result = op.execute(Literal("hallo", lang="de")) + assert _is_simple_literal(result) + assert str(result) == "hallo" + + def test_typed_literal_yields_lexical_form(self, settings): + # §4.2: STR(42) → "42" (simple literal) op = Operation.get("Str")(settings=settings) - bn = BNode("b1") - result = op.execute(bn) - assert isinstance(result, Literal) - assert str(result) == str(bn) + result = op.execute(Literal(42)) + assert _is_simple_literal(result) + assert str(result) == "42" + + def test_bnode_raises_type_error(self, settings): + # §4.2: STR accepts a literal or an IRI; a blank node is a type error + op = Operation.get("Str")(settings=settings) + with pytest.raises(TypeError): + op.execute(BNode("b1")) def test_non_term_raises_type_error(self, settings): - # Strict Type Checking property: "TypeError raised for mismatched input types" + # §3.7 strict typing op = Operation.get("Str")(settings=settings) with pytest.raises(TypeError): op.execute([1, 2, 3]) - @pytest.mark.skip(reason="UNCLEAR(spec): result Literal datatype (xsd:string vs simple literal) unspecified") - def test_result_datatype(self, settings): - pass - class TestStrJson: def test_string_input_via_json(self, settings): - # JSON arg key derived from existing fixture tests/fixtures/positive/simple-operation.json + # §4.2 JSON: input: URI + Literal op = Operation.get("Str")(settings=settings) result = op.execute_json({"input": "hello"}) - assert isinstance(result, Literal) + assert _is_simple_literal(result) assert str(result) == "hello" def test_nested_uri_op(self, settings): @@ -56,5 +76,5 @@ def test_nested_uri_op(self, settings): result = op.execute_json( {"input": {"@op": "URI", "args": {"input": "http://example.org/x"}}} ) - assert isinstance(result, Literal) + assert _is_simple_literal(result) assert str(result) == "http://example.org/x" diff --git a/tests/unit/test_struuid.py b/tests/unit/test_struuid.py index 899624f..ed52dce 100644 --- a/tests/unit/test_struuid.py +++ b/tests/unit/test_struuid.py @@ -5,7 +5,6 @@ from __future__ import annotations -import pytest from rdflib import Literal from web_algebra.operation import Operation @@ -23,9 +22,18 @@ def test_two_calls_differ(self, settings): b = op.execute() assert str(a) != str(b) - @pytest.mark.skip(reason="UNCLEAR(spec): UUID format (UUID4? hyphenated? case?) not specified") def test_uuid_format(self, settings): - pass + # §4.2: simple literal per `simple literal STRUUID()`, holding an + # RFC 4122 version-4 UUID in lowercase hyphenated form + import re + + op = Operation.get("STRUUID")(settings=settings) + result = op.execute() + assert result.datatype is None and result.language is None + assert re.fullmatch( + r"[0-9a-f]{8}-[0-9a-f]{4}-4[0-9a-f]{3}-[89ab][0-9a-f]{3}-[0-9a-f]{12}", + str(result), + ) class TestSTRUUIDJson: diff --git a/tests/unit/test_substitute.py b/tests/unit/test_substitute.py index 2725d56..a33d3b4 100644 --- a/tests/unit/test_substitute.py +++ b/tests/unit/test_substitute.py @@ -1,12 +1,15 @@ -"""Spec: formal-semantics.md "Substitute - Replace variables in SPARQL queries" -Abstract: Literal × Literal × Term → Literal -Python: def execute(self, query: Literal, var: Literal, binding_value: Any) -> Literal +"""Spec: formal-semantics.md §4.3 "Substitute — textually substitute one +SPARQL variable with a Term" +Abstract: Literal × Literal × (URI + Literal) → Literal +- Matches both `?var` and `$var` at token boundaries. +- URI serializes as ``; Literal as a quoted literal with its language + tag or datatype. BNode raises TypeError. """ from __future__ import annotations import pytest -from rdflib import Literal, URIRef +from rdflib import BNode, Literal, URIRef from web_algebra.operation import Operation @@ -27,12 +30,64 @@ def test_wrong_query_type_raises(self, settings): with pytest.raises(TypeError): op.execute(URIRef("not-a-query"), Literal("x"), URIRef("http://example.org/foo")) - @pytest.mark.skip(reason="UNCLEAR(spec): SPARQL variable syntax — `?var`, `$var`, or both? How are URIRef/Literal binding values serialized into the query?") - def test_replacement_form(self, settings): - pass + def test_question_mark_variable_replaced_with_iri(self, settings): + # §4.3: `?var` matched; URI serializes as + op = Operation.get("Substitute")(settings=settings) + result = op.execute( + Literal("DESCRIBE ?x"), Literal("x"), URIRef("http://example.org/foo") + ) + assert "" in str(result) + assert "?x" not in str(result) + + def test_dollar_variable_replaced(self, settings): + # §4.3: `$var` matched too + op = Operation.get("Substitute")(settings=settings) + result = op.execute( + Literal("DESCRIBE $x"), Literal("x"), URIRef("http://example.org/foo") + ) + assert "" in str(result) + assert "$x" not in str(result) + + def test_token_boundary_respected(self, settings): + # §4.3: matches at token boundaries — ?xy must not be touched when + # substituting ?x + op = Operation.get("Substitute")(settings=settings) + result = op.execute( + Literal("SELECT ?xy WHERE { ?x ?p ?xy }"), + Literal("x"), + URIRef("http://example.org/foo"), + ) + assert "?xy" in str(result) + + def test_literal_value_serialized_with_datatype(self, settings): + # §4.3: Literal serializes as a quoted literal with its datatype + op = Operation.get("Substitute")(settings=settings) + result = op.execute( + Literal("SELECT * WHERE { ?s ?p ?x }"), + Literal("x"), + Literal("42", datatype=URIRef("http://www.w3.org/2001/XMLSchema#integer")), + ) + assert '"42"' in str(result) + assert "http://www.w3.org/2001/XMLSchema#integer" in str(result) + + def test_bnode_value_raises_type_error(self, settings): + # §4.3: a blank-node label in a query is a fresh variable, not a + # reference — BNode raises TypeError + op = Operation.get("Substitute")(settings=settings) + with pytest.raises(TypeError): + op.execute(Literal("DESCRIBE ?x"), Literal("x"), BNode("b1")) class TestSubstituteJson: - @pytest.mark.skip(reason="UNCLEAR(spec): JSON arg keys for Substitute not given by spec or existing fixtures") def test_json_dispatch(self, settings): - pass + # §4.3 JSON: query · var · binding (Term or SPARQL JSON term form §2.4) + op = Operation.get("Substitute")(settings=settings) + result = op.execute_json( + { + "query": "DESCRIBE ?x", + "var": "x", + "binding": {"type": "uri", "value": "http://example.org/foo"}, + } + ) + assert isinstance(result, Literal) + assert "" in str(result) diff --git a/tests/unit/test_uri.py b/tests/unit/test_uri.py index 9ceceea..a25eeda 100644 --- a/tests/unit/test_uri.py +++ b/tests/unit/test_uri.py @@ -1,7 +1,8 @@ -"""Spec: formal-semantics.md "URI - Convert term to URI reference" -Abstract: Term → URI -Python: def execute(self, term: rdflib.term.Node) -> rdflib.URIRef -Plus the Strict Type Checking property (lines 291-295). +"""Spec: formal-semantics.md §4.2 "URI — cast a Term to a URI" +Abstract: (URI + Literal) → URI +- URI input returned as-is; Literal yields the URI of its lexical form. +- BNode raises TypeError (a blank node has no IRI). +- The lexical form is NOT validated against RFC 3986. """ from __future__ import annotations @@ -31,15 +32,19 @@ def test_non_term_raises_type_error(self, settings): with pytest.raises(TypeError): op.execute(42) - @pytest.mark.skip(reason="UNCLEAR(spec): URI(BNode) — spec lists BNode as a Term but doesn't define this case") - def test_bnode_input(self, settings): + def test_bnode_raises_type_error(self, settings): + # §4.2: a BNode raises TypeError — a blank node has no IRI op = Operation.get("URI")(settings=settings) - op.execute(BNode("b1")) + with pytest.raises(TypeError): + op.execute(BNode("b1")) - @pytest.mark.skip(reason="UNCLEAR(spec): URI(Literal whose lexical form is not a valid URI) unspecified") - def test_invalid_uri_literal(self, settings): + def test_invalid_uri_literal_is_not_validated(self, settings): + # §4.2: the lexical form is not validated against RFC 3986 — + # garbage in, garbage out op = Operation.get("URI")(settings=settings) - op.execute(Literal("not a uri")) + result = op.execute(Literal("not a uri")) + assert isinstance(result, URIRef) + assert str(result) == "not a uri" class TestURIJson: diff --git a/tests/unit/test_value.py b/tests/unit/test_value.py index 3696a25..e7f5ef1 100644 --- a/tests/unit/test_value.py +++ b/tests/unit/test_value.py @@ -1,7 +1,12 @@ -"""Spec: formal-semantics.md "Value - Access variables and context values" -Abstract: String × Context × VariableStack → Any -Python: def execute(self, name: str, context: Any, variable_stack: List[Dict[str, Any]]) -> Any -Plus Variable System property (lines 308-311) and Context System property (lines 314-318). +"""Spec: formal-semantics.md §4.1 "Value" with §3.4 (variable environment) +and §3.5 (focus). +Abstract: String → Any +- `$name` searches variable scopes innermost to outermost; miss → ValueError. +- Unprefixed `name` looks up in the focus item; the item shapes are closed: + Binding → bound term, mapping → member value; a miss or any other item + shape → ValueError. +- The `$` sigil decides the lookup domain, so the two never shadow each other. +- Returns the value as-is (xsl:sequence semantics, no string conversion). """ from __future__ import annotations @@ -26,20 +31,58 @@ def test_lookup_falls_back_to_outer_scope(self, settings): result = op.execute("$x", {}, stack) assert result == Literal("outer") - @pytest.mark.skip(reason="UNCLEAR(spec): which context container shapes does Value support? Context type is `Any` (line 315) and the narrative names ResultRow but doesn't enumerate other shapes (dict? attribute-bearing object? both?)") - def test_context_lookup(self, settings): - pass + def test_mapping_context_lookup(self, settings): + # §3.5: mapping focus item → member value + op = Operation.get("Value")(settings=settings) + result = op.execute("city", {"city": Literal("Vilnius")}, []) + assert result == Literal("Vilnius") + + def test_lookup_unwraps_the_focus(self, settings): + # §3.5: the unprefixed lookup targets the focus *item* + from web_algebra.focus import Focus + + op = Operation.get("Value")(settings=settings) + focus = Focus(item={"city": Literal("Vilnius")}, position=1, size=1) + assert op.execute("city", focus, []) == Literal("Vilnius") + + def test_unsupported_item_shape_raises_value_error(self, settings): + # §3.5: the item shapes are closed (Binding + mapping) — anything + # else raises ValueError, even if the host object happens to carry + # an attribute of that name + class Item: + city = Literal("Kaunas") + + op = Operation.get("Value")(settings=settings) + with pytest.raises(ValueError): + op.execute("city", Item(), []) + + def test_sigil_selects_lookup_domain(self, settings): + # §3.4: `$name` reads the variable stack, plain `name` the context — + # the same name in both never shadows. + op = Operation.get("Value")(settings=settings) + context = {"x": Literal("from-context")} + stack = [{"x": Literal("from-stack")}] + assert op.execute("$x", context, stack) == Literal("from-stack") + assert op.execute("x", context, stack) == Literal("from-context") - @pytest.mark.skip(reason="UNCLEAR(spec): precedence when same name appears in both context and stack") - def test_context_vs_stack_precedence(self, settings): - pass + def test_missing_variable_raises_value_error(self, settings): + # §3.7: unknown variable in `$name` lookup → ValueError + op = Operation.get("Value")(settings=settings) + with pytest.raises(ValueError): + op.execute("$missing", {}, []) - @pytest.mark.skip(reason="UNCLEAR(spec): behavior on missing name — error class unspecified") - def test_missing_name(self, settings): - pass + def test_missing_context_member_raises_value_error(self, settings): + # §3.7: context lookup miss → ValueError + op = Operation.get("Value")(settings=settings) + with pytest.raises(ValueError): + op.execute("missing", {"other": Literal("v")}, []) class TestValueJson: - @pytest.mark.skip(reason="UNCLEAR(spec): JSON arg key for Value not given by spec or existing fixtures") def test_json_dispatch(self, settings): - pass + # §4.1 JSON: name: String (plain JSON string, `$` prefix for variables) + op = Operation.get("Value")(settings=settings) + result = op.execute_json( + {"name": "$x"}, [{"x": Literal("bound")}] + ) + assert result == Literal("bound") diff --git a/tests/unit/test_variable.py b/tests/unit/test_variable.py index 8df6244..f4b5a94 100644 --- a/tests/unit/test_variable.py +++ b/tests/unit/test_variable.py @@ -1,13 +1,14 @@ -"""Spec: formal-semantics.md "Variable - Set variables in current scope (XSLT-style)" -Abstract: String × Any × VariableStack → ⊥ -Python: def execute(self, name: str, value: Any, variable_stack: List[Dict[str, Any]]) -> None -Plus Variable System property (lines 308-311). +"""Spec: formal-semantics.md §4.1 "Variable — bind a name in the current scope" +Abstract: String × Any → Unit +- Binds in the innermost scope; rebinding the same name overwrites (§3.4). +- The JSON layer returns Unit (None). +- Scope creation belongs to sequences and ForEach iterations, not to Variable. """ from __future__ import annotations -import pytest from rdflib import Literal +from rdflib.namespace import XSD from web_algebra.operation import Operation @@ -22,16 +23,39 @@ def test_binds_into_current_scope(self, settings): result = value_op.execute("$x", {}, stack) assert result == Literal("v") - @pytest.mark.skip(reason="UNCLEAR(spec): `⊥` (bottom) return type — what does execute_json return on the JSON layer?") - def test_return_value(self, settings): - pass - - @pytest.mark.skip(reason="UNCLEAR(spec): line 311 self-contradiction — does Variable push a new scope or write into the current one?") - def test_scope_management(self, settings): - pass + def test_returns_unit(self, settings): + # §4.1: Abstract String × Any → Unit; JSON layer returns None + op = Operation.get("Variable")(settings=settings) + assert op.execute("x", Literal("v"), [{}]) is None + + def test_binds_into_innermost_scope_only(self, settings): + # §3.4: Variable binds in the innermost scope; it does not push or + # pop scopes itself. + op = Operation.get("Variable")(settings=settings) + stack = [{}, {}] + op.execute("x", Literal("v"), stack) + assert len(stack) == 2 + assert "x" not in stack[0] + assert stack[1]["x"] == Literal("v") + + def test_rebinding_overwrites(self, settings): + # §3.4: rebinding a name in the same scope overwrites it + op = Operation.get("Variable")(settings=settings) + stack = [{}] + op.execute("x", Literal("first"), stack) + op.execute("x", Literal("second"), stack) + assert stack[0]["x"] == Literal("second") class TestVariableJson: - @pytest.mark.skip(reason="UNCLEAR(spec): JSON arg keys for Variable not given by spec or existing fixtures") def test_json_dispatch(self, settings): - pass + # §4.1 JSON: name: String · value: any form; returns Unit (None) + op = Operation.get("Variable")(settings=settings) + stack = [{}] + result = op.execute_json({"name": "x", "value": "v"}, stack) + assert result is None + value_op = Operation.get("Value")(settings=settings) + # §2.2: the scalar "v" coerces to an xsd:string Literal + assert value_op.execute_json({"name": "$x"}, stack) == Literal( + "v", datatype=XSD.string + ) diff --git a/tests/unit/test_write_contract.py b/tests/unit/test_write_contract.py new file mode 100644 index 0000000..8bd5971 --- /dev/null +++ b/tests/unit/test_write_contract.py @@ -0,0 +1,288 @@ +"""Spec: formal-semantics.md §4.4 — the contract shared by the writes POST, PUT +and PATCH, and §3.7's error table. +- A write returns a single-row Result with variables `status` (xsd:integer + HTTP status) and `url`: the response's Location when it has one, otherwise + the effective request URI after redirects. +- A response outside 2xx is an error (ValueError, not urllib's HTTPError); + transport failures (no response) propagate unwrapped. +- If-Match: the entity tag is read with HEAD, sending the same Accept as the + write, and sent as If-Match; a resource that does not exist, or has no tag, + is written unconditionally. +- `Filter(Filter(PUT(…), 1), "url")` is the URI a write reports (§4.1). + +HTTP is stubbed at the urllib boundary (tests/http_stub.py); no network. +""" + +from __future__ import annotations + +import urllib.error + +import pytest +from rdflib import Graph, Literal, URIRef +from rdflib.namespace import XSD +from rdflib.query import Result + +from tests.http_stub import StubResponse, default_handler, request_header +from web_algebra.operation import Operation + +EX = "http://example.org/" +DOC = EX + "doc" + +WRITES = ["POST", "PUT", "PATCH"] + + +def _graph() -> Graph: + g = Graph() + g.add((URIRef(DOC), URIRef(EX + "p"), Literal("v"))) + return g + + +def _write(method: str, settings, url: str = DOC): + op = Operation.get(method)(settings=settings) + if method == "PATCH": + return op.execute( + URIRef(url), Literal(f"INSERT DATA {{ <{DOC}> <{EX}p> \"v\" }}") + ) + return op.execute(URIRef(url), _graph()) + + +def _writes_answered(method: str, response_factory, head=None): + """A handler answering `method` with `response_factory()` and HEAD with + `head()` (default: 200, no ETag).""" + + def handler(request): + m = request.get_method() + if m == method: + return response_factory() + if m == "HEAD": + return head() if head else StubResponse(200) + return default_handler(request) + + return handler + + +def _single_row(result): + assert isinstance(result, Result) + rows = list(result) + assert len(rows) == 1 + return rows[0] + + +@pytest.mark.parametrize("method", WRITES) +class TestWriteResult: + def test_returns_single_row_result_with_status_and_url( + self, method, settings, http_stub + ): + # §4.4: a single-row Result, variables status (xsd:integer) and url + http_stub.handler = _writes_answered(method, lambda: StubResponse(200)) + row = _single_row(_write(method, settings)) + status = row["status"] + assert isinstance(status, Literal) + assert status.datatype == XSD.integer + assert int(status) == 200 + assert str(row["url"]) == DOC + + def test_status_is_the_response_status(self, method, settings, http_stub): + http_stub.handler = _writes_answered(method, lambda: StubResponse(204)) + row = _single_row(_write(method, settings)) + assert int(row["status"]) == 204 + + def test_url_is_location_when_present(self, method, settings, http_stub): + # §4.4: url is the response's Location when it has one, as when a + # POST to a container creates a child + child = EX + "container/child" + http_stub.handler = _writes_answered( + method, lambda: StubResponse(201, {"Location": child}) + ) + row = _single_row(_write(method, settings, EX + "container/")) + assert str(row["url"]) == child + + def test_url_is_effective_request_uri_after_redirects( + self, method, settings, http_stub + ): + # §4.4: without Location, url is the effective request URI after + # redirects (what urllib reports as the response URL) + moved = EX + "moved" + http_stub.handler = _writes_answered( + method, lambda: StubResponse(200, url=moved) + ) + row = _single_row(_write(method, settings)) + assert str(row["url"]) == moved + + def test_relative_location(self, method, settings, http_stub): + # §4.4: a relative Location is resolved against the effective request + # URI (RFC 3986 §5) + http_stub.handler = _writes_answered( + method, lambda: StubResponse(201, {"Location": "child"}) + ) + row = _single_row(_write(method, settings, EX + "container/")) + assert str(row["url"]) == EX + "container/child" + + def test_relative_location_after_redirect(self, method, settings, http_stub): + # §4.4: the base is the *effective* request URI, after redirects + http_stub.handler = _writes_answered( + method, + lambda: StubResponse(201, {"Location": "child"}, url=EX + "moved/"), + ) + row = _single_row(_write(method, settings, EX + "container/")) + assert str(row["url"]) == EX + "moved/child" + + +@pytest.mark.parametrize("method", WRITES) +class TestWriteErrors: + @pytest.mark.parametrize("status", [400, 403, 404, 409, 412, 428, 500]) + def test_non_2xx_raises_value_error_not_http_error( + self, method, status, settings, http_stub + ): + # §4.4/§3.7: a write answered outside 2xx → ValueError; the + # transport's HTTPError is not what surfaces + http_stub.handler = _writes_answered(method, lambda: StubResponse(status)) + with pytest.raises(ValueError) as exc_info: + _write(method, settings) + assert not isinstance(exc_info.value, urllib.error.HTTPError) + + def test_transport_failure_propagates_unwrapped(self, method, settings, http_stub): + # §3.7: no response → URLError, unwrapped + def handler(request): + if request.get_method() == method: + raise urllib.error.URLError("connection refused") + if request.get_method() == "HEAD": + return StubResponse(404) + return default_handler(request) + + http_stub.handler = handler + with pytest.raises(urllib.error.URLError) as exc_info: + _write(method, settings) + assert not isinstance(exc_info.value, ValueError) + + +@pytest.mark.parametrize("method", WRITES) +class TestWriteIfMatch: + def test_etag_from_head_is_sent_as_if_match(self, method, settings, http_stub): + # §4.4: the entity tag is read with HEAD and sent as If-Match + http_stub.handler = _writes_answered( + method, + lambda: StubResponse(200), + head=lambda: StubResponse(200, {"ETag": '"abc123"'}), + ) + _write(method, settings) + writes = http_stub.with_method(method) + assert len(writes) == 1 + assert request_header(writes[0], "If-Match") == '"abc123"' + + def test_head_precedes_write_on_the_same_uri_with_same_accept( + self, method, settings, http_stub + ): + # §4.4: HEAD on the resource, with the same Accept as the write (the + # tag names a negotiated variant), immediately before the write + http_stub.handler = _writes_answered( + method, + lambda: StubResponse(200), + head=lambda: StubResponse(200, {"ETag": '"abc123"'}), + ) + _write(method, settings) + methods = http_stub.methods() + assert "HEAD" in methods + head_index = methods.index("HEAD") + write_index = methods.index(method) + assert head_index < write_index + head = http_stub.requests[head_index] + write = http_stub.requests[write_index] + assert head.full_url == write.full_url == DOC + assert request_header(head, "Accept") == request_header(write, "Accept") + + def test_no_if_match_when_head_fails(self, method, settings, http_stub): + # §4.4: a resource that does not exist is written unconditionally + http_stub.handler = _writes_answered( + method, lambda: StubResponse(201), head=lambda: StubResponse(404) + ) + _write(method, settings) + writes = http_stub.with_method(method) + assert len(writes) == 1 + assert request_header(writes[0], "If-Match") is None + + def test_no_if_match_when_head_has_no_etag(self, method, settings, http_stub): + # §4.4: a resource that has no tag is written unconditionally + http_stub.handler = _writes_answered( + method, lambda: StubResponse(200), head=lambda: StubResponse(200) + ) + _write(method, settings) + writes = http_stub.with_method(method) + assert len(writes) == 1 + assert request_header(writes[0], "If-Match") is None + + def test_412_after_head_is_an_error(self, method, settings, http_stub): + # §4.4: a write made by another client between HEAD and the write is + # answered 412, an error like any other non-2xx answer + http_stub.handler = _writes_answered( + method, + lambda: StubResponse(412), + head=lambda: StubResponse(200, {"ETag": '"old"'}), + ) + with pytest.raises(ValueError): + _write(method, settings) + + +class TestReportedUrlLookup: + def test_filter_filter_put_url_is_the_reported_uri(self, settings, http_stub): + # §4.4/§4.1: Filter(Filter(PUT(…), 1), "url") is the URI a write + # reports + result = Operation.process_json( + settings, + { + "@op": "Filter", + "args": { + "input": { + "@op": "Filter", + "args": { + "input": { + "@op": "PUT", + "args": { + "url": {"@id": DOC}, + "data": {"@id": DOC, EX + "p": "v"}, + }, + }, + "expression": 1, + }, + }, + "expression": "url", + }, + }, + ) + assert result == URIRef(DOC) + + def test_for_each_over_write_makes_the_row_the_focus(self, settings, http_stub): + # §4.4: ForEach(select: PUT(…), operation: GET(url: Value(url))) + # dereferences the written document + def handler(request): + if request.get_method() == "GET": + return StubResponse( + 200, + {"Content-Type": "text/turtle"}, + f"<{DOC}> <{EX}p> \"v\" .".encode(), + ) + return default_handler(request) + + http_stub.handler = handler + result = Operation.process_json( + settings, + { + "@op": "ForEach", + "args": { + "select": { + "@op": "PUT", + "args": { + "url": {"@id": DOC}, + "data": {"@id": DOC, EX + "p": "v"}, + }, + }, + "operation": { + "@op": "GET", + "args": {"url": {"@op": "Value", "args": {"name": "url"}}}, + }, + }, + }, + ) + assert len(result) == 1 + assert isinstance(result[0], Graph) + assert [r.full_url for r in http_stub.with_method("GET")] == [DOC] diff --git a/uv.lock b/uv.lock index 72cdf49..a41e218 100644 --- a/uv.lock +++ b/uv.lock @@ -888,7 +888,7 @@ wheels = [ [[package]] name = "web-algebra" -version = "1.4.0" +version = "2.0.0" source = { editable = "." } dependencies = [ { name = "mcp", extra = ["cli"] },