Skip to content

Clausal Prolog — Seam Syntax Design

This page documents the seam (.seam): Clausal Prolog's Python-syntax surface. See Clausal Prolog for the ISO-syntax .clausal surface.

The seam uses Python's parser, and so it conforms to Python's grammar. Like Prolog, it describes Horn Clauses - simple constructs that facilitate the representation of facts and rules, backed by strong mathematical formalisms, to facilitate high-level reasoning with useful guarantees.

The seam uses 'grammatical holes' which are and must remain syntactically valid, but have no semantic purpose in Python, and are therefore never used in practice. The seam uses a very small set of these 'holes' to allow the free mixing of logic code with Python code. Python and logic programming code in the same file means better code cohesion, easing development.

Quick navigation: Variables · Constants · Escape operators · Unification · Clauses · Lists · Arithmetic · Constraints · DCGs · Lambdas · Meta-predicates · Cheatsheet


Trailing comma rule

Expression statements ending in , are interpreted as logical terms (facts or goals).

In standard Python, an expression statement that ends in a comma produces a tuple which is then discarded. So this is never used despite being valid (a grammatical hole). It also means these logic terms are easy to cut, copy, paste, and indent — unlike Prolog, they don't require . terminators.


Escape operators

Three double-prefix operators demarcate the boundary between Python and logic code. They must not be split with spaces (e.g. - - would be parsed as double negation).

Operator Meaning
--term In Python code: build a logic term (a cell) — and, in goal position (if --g:, for X in --g:), run it as a query. See Python integration
++expr In logic code: a Python value — the expression is evaluated and passed in unconverted (in a clause body, at search time)
~~expr Capture the expression as a Python ast node without running it (works anywhere)
--X inside a thunk The logic variable X — see Marking a variable inside a thunk

-- was chosen because: - it doesn't introduce a new keyword or clobber any identifier - double negation is rare in Python code (and if you really want it, type it as '- -x') - it is visually prominent and quick to type

Example:

# Logic term containing a Python value:
point(++x_coord, ++y_coord),

# Python expression containing a logic term (a cell):
my_term = --foo(1, bar(2))          # ('foo', 1, ('bar', 2))

# Capture Python code as AST without running it:
ast_node = ~~(x + y * z)


Logic variables

Three conventions are recognised:

Leading single underscore — any identifier whose first character is _, excluding dunders (__) and the bare anonymous variable _:

_x, _head, _rest   # logic variables (leading-underscore style)

ALL_CAPS — any identifier where every cased character is uppercase and there is at least one cased character (underscores and digits are allowed inside):

X, HEAD, REST, N1, MAX_OF   # logic variables (ALL-CAPS style)

TitleCase — an initial capital followed by at least one lowercase letter:

Foo, Total, FooBar   # logic variables (TitleCase style, the ISO Prolog spelling)

All three styles may be used in the same file, and a name in any of them is a logic variable wherever a value goes. ALL_CAPS is the preferred style for new code; leading-underscore is available when a lowercase variable name is desired; TitleCase is the ISO Prolog spelling and is accepted so that Prolog can be read and written without transliteration.

One asymmetry: TitleCase in functor position

The three styles are not interchangeable in every position. In functor position — a clause head's functor, or a goal called in a body — a bare TitleCase name is a load-time error:

p(Foo) <- (bar(Foo))          # fine: Foo is a variable in term position
Foo(X) <- (bar(X))            # SyntaxError: Foo is TitleCase
p(X) <- (bar(Fraction(1, 3))) # SyntaxError: reach the class as ++Fraction

The reason is that a variable in functor position does not mean call/N in the seam — it is the unit-literal sugar, so X(newton) builds a quantity rather than calling X. If TitleCase were a variable there too, a bare Python class in a clause body would stop being a clear load error and quietly become a units expression that fails much later. So FOO(3) remains legal units sugar with a computed unit, while Foo(3) is refused.

The remedy the error names is the ++ escape: a Python class is reached as ++ClassName. A name bound by an -import_from list, and the injected Undefined, keep their binding in every position and are not affected.

The refusal is about the bare spelling only. A functor is named by an atom, and a single-quoted token is an atom whatever its capitalisation — so quoting says "the atom Foo", not "the variable Foo", and every functor position accepts it:

'Foo'(1),                     # the fact Foo(1) — Foo/1, not a variable
'Foo'(X) <- (bar(X))          # a rule head for the same predicate
p(X) <- ('Foo'(X))            # and a goal that calls it

This is ISO's own asymmetry: 'Foo'(1) is a compound term there too, while the bare Foo(1) is a syntax_error(variable_cannot_be_functor). As everywhere else, a double-quoted literal is not an atom spelling and never names a functor (ISO 6.3.3).

A declaration list takes the quoted spelling too, which is how such a predicate is exported — -module(m, ['Foo'(X)]), or the ISO arity form -module(m, ['Foo'/1]) — and -private reads it the same way.

One seam limit applies to anything that names a predicate — a head or a declaration entry — and not to goals: the name must spell a plain, non-keyword name. 'foo bar'(X) and 'not'(X) may be called; they may not be defined or declared.

A single _ is the anonymous variable — it never stores a value, and unification against it always succeeds (matching Python's and Prolog's existing convention).

Nothing is carved out of these conventions. Until 2026-09-11 one shape was: an identifier with exactly one leading and one trailing underscore (_PI_) was a constant, never a variable. Constants are spelled like atoms now (see Constants), so _PI_ is an ordinary leading-underscore variable like _x, and the rule is exactly two clauses with no exceptions.

Logic variables are not declared; they come into existence by appearing in logical context. They work differently from Python variables: they can be unbound, and their bindings are undone on backtracking. This difference warrants a clear visual marker.

The seam originally departed from Prolog here, reserving TitleCase for atoms and functors on the grounds that in Foo(Bar) it would be ambiguous whether Bar was an atom or a variable. That worked example is now resolved the other way, and the ambiguity it feared does not arise: position decides, not spelling. In Foo(Bar) the argument Bar is a logic variable, and the functor Foo is refused outright — so there is no shape in which one TitleCase name could be read two ways.

Why the extra style is worth having: - It is the ISO Prolog spelling, so Prolog sources read and translate without transliterating every variable. See Prolog Translation for the full mapping. - ALL-CAPS remains available and is still the preferred style for new seam code: single letters like X, Y, N are universally understood as logic variables from mathematics, and ALL-CAPS marks the variable role unmistakably to a reader who also writes Python. - Leading underscore (_x) is available when a lowercase-looking variable name reads better.

The one cost is that a TitleCase name no longer looks like a Python class to a reader skimming a clause. The functor rule above is what keeps that from becoming a silent error: the position where a class name would actually be used is exactly the position that still refuses it.

Singleton variables and _UNUSED

A named variable that occurs exactly once in its clause binds nothing — almost always a typo (a dropped letter, a copy-paste that missed one occurrence). The seam loader warns on this by default, in both variable styles:

test("bug: wrong var name") <- (
    length([1, 2, 3], N),
    N1 == N + 1     # meant N, typo'd N1 — N1 is a singleton
)
ClausalSingletonWarning: m.seam:1: singleton variable `N1` — a variable occurring once
binds nothing. Misspelling? Rename to `N1_UNUSED` (or `_`) if deliberate, or add
-allow_singletons to the file

(The reported line is where the clause statement starts, not the specific goal inside it — useful context for a multi-goal body.)

If a single-occurrence variable is deliberate — you want its name for readability but never read its value — suppress the warning with the _UNUSED suffix. This is the sole canonical spelling: _UNUSED (uppercase, exact) is recognised; _unused, _Unused, and every other casing are not exempt — one spelling, one grep target, no guessing which files opted out.

handle(EVENT, REASON_UNUSED) <- (EVENT == 'click')   # REASON_UNUSED never read — fine

If you don't need the name at all, prefer the bare anonymous variable _ — it never warns, in any number of occurrences per clause.

The check runs per clause, so _UNUSED on a genuinely-reused name is a real bug the lint catches too — this is the inverse lint, modeled on SWI-Prolog's _X warning:

ClausalSingletonWarning: m.seam:1: variable `N1_UNUSED` is marked _UNUSED but occurs more
than once in its clause

To opt an entire file out — a fixture that deliberately demonstrates the pattern, for example — use -allow_singletons:

-allow_singletons

test("most general query") <- var(SOME_UNBOUND_VAR)

Chained comparisons count their middle operand twice

1 < X < 10 desugars to two goals sharing X (1 < X and X < 10), so X counts as 2 occurrences even though it appears once in the source. A variable that occurs only as the middle term of a chained comparison therefore never warns as a singleton — and, by the same token, renaming it to X_UNUSED would trip the inverse lint (marked _UNUSED but occurs more than once), since the desugaring still duplicates it.

Known gap: DCG/EDCG bodies are not yet linted

>> grammar rules (both plain DCGs and EDCGs) do not run through the singleton check — a genuine singleton inside a DCG body currently goes unwarned. Ordinary <- clauses (including the ones a DCG rule rewrites to, if you write them by hand) are fully covered.


Constants

area == pi * R**2 instead of (is_pi(PI), area == PI * R**2) or a Python math.pi escape: a constant is a module-level name, bound once to a ground value at load time, that reads exactly like an ordinary term argument.

Lexical rule — a constant is spelled like an atom. That is stated as the complement of the logic-variable rule: any identifier that rule does not claim, which is to say one that is neither underscore-led nor capital-initial.

pi, max_retries, 円周率     # constant names
_pi, Pi, PI                  # rejected — all three are logic variables

円周率 is in that list on purpose. An uncased script has no lowercase form and no capital one, so a rule phrased as "lowercase" would refuse it while offering nothing in its place; a rule phrased as "not a variable" admits it, because it is not a variable.

This spelling changed on 2026-09-11

A constant used to be written _PI_ — one leading and one trailing underscore. To any Prolog reader a leading underscore marks a variable, so the old spelling read as the one thing a constant is not, and it cost the variable rule an exception in five places. The old spelling is refused with a message naming the replacement.

Reaching the value. Because the name is atom-shaped, a constant lives in the same namespace as an atom, and the two ways of writing it differ in when the value is read:

written means bound
constant(pi) the value, looked up in module globals when the goal runs
pi the atom 'pi' —

constant(pi) is the retrieval form. The older ++pi still works — a constant is a module global and ++ is the Python escape — but it says "what follows is Python", which is the one thing a constant reference is not. The parentheses delimit the name, and what goes inside must be a single atom, so there is no shape here that could be read as a Python expression. An undeclared name is a load-time error, where the escape deferred to a NameError when the goal eventually ran.

constant(pi) is the spelling to reach for. It is what makes a constant late-bound, so that changing a declaration changes every use; and it is explicit, which matters when the same file also uses the atom. A name may be both a constant and an atom: listing it in -module/-private declares the atom (so a bare max_retries is allowed under strict atoms) and leaves the constant in place, so constant(max_retries) is still 3.

Declaring

-constant_value(pi, 3.14159)
-constant_value(max_retries, 3)

area(R, AREA) <- (AREA == ++pi * R**2)

test("area of radius 2") <- (
    area(2, AREA),
    AREA == 12.56636
)

One constant per directive, positionally. A later -constant_value may reference a constant an earlier one declared. The right-hand side accepts:

  • scalar literals (3.14159, "eur", True),
  • a previously-declared constant (`-constant_value(base, 10)
  • a previously-declared constant (-constant_value(limit, base * 4 + 2)),
  • a declared atom,
  • unary/binary arithmetic over those,
  • a ++(expr) escape, evaluated as raw Python at load time, and
  • structured literals — lists, tuples, sets, dicts, and functor calls, nested arbitrarily, mixing any of the above at any depth:
-constant_value(pi, ++__import__('math').pi)

++ RHS values can be machine-dependent

++(expr) is evaluated once, when the file loads — nothing stops it from calling something that isn't reproducible across machines or runs: -constant_value(n_workers, ++os.cpu_count()) is legal, and will bind a different value on a different machine. This is a documented caveat, not a guardrail — if reproducibility matters, don't reach for ++ in a -constant_value RHS.

Structured constants

A structured RHS builds a real logic term — the same term the identical literal would build in a clause body, with the same unification semantics — not a Python value:

-module(m, [point(X, Y)])
-private([mn, mx, red, green])

-constant_value(country_codes, ['au', 'al', 'za'])
-constant_value(origin, point(0, 0))
-constant_value(limits, {mn: 1, mx: 99})
-constant_value(flags, {red, green})

Lists, tuples, sets, and dicts lower through the same construction a clause body's literal of the same shape uses (a list is a plain list, a set is a SetTerm, a dict is a DictTerm), and a functor call (point(0, 0)) constructs a real instance of that functor's class — so a structured constant unifies exactly as the equivalent literal would, indexes the same way, and carries no extra runtime cost per reference.

A functor used in a structured RHS must already be declared above the -constant_value directive — via -module, -private, -dynamic, an earlier RULE clause (point(X, Y) <- (...)) defining it, an earlier properly-terminated bodyless FACT (point(1, 2), — the trailing comma is what makes it a fact at all; point(1, 2) with no comma is a bare, unregistered expression statement, not a declaration — see the comma-optional fact rule above), or imported via -import_from/-import_module. This is the same source-order rule that makes the functor's class statement execute before the constant's assignment does; an undeclared functor is a located, load-time SyntaxError naming the remedy:

SyntaxError: -constants: `p` RHS calls `point(...)`, which is not a declared functor above
this -constant_value directive — declare it with -module/-private/-dynamic first, or
import it with -import_from/-import_module

Every declaration must still be fully ground — no unbound logic variable may appear anywhere in the computed value, structured RHS included: -constant_value(l, [1, X, 3]) is a located, load-time SyntaxError, not a freshly-minted Var. (An unground value that only a ++() escape could produce still hits the runtime ConstantNotGroundError backstop, before any clause compiles.) A ++() escape is legal as an element inside a structured RHS ([1, ++(2 + 3), 3]) — it lowers exactly as it would at the top level.

Structured constants are frozen, not hidden

A structured constant's value is bound to the module global once, at load time, and that value is immutable: a list constant is a frozen list, a dict constant's backing store is a frozen dict, and a ++()-escape-built raw Python set is a frozen set (a source-level {...} set literal lowers to SetTerm, which is already immutable, nothing further to freeze; a ++()-escape-built frozenset is likewise already immutable and is returned as-is, not re-wrapped). Any mutating call reached from Python — .append, __setitem__, .add, |=, and the like, most commonly reached through a ++() escape holding a reference to the constant — raises TypeError rather than silently corrupting the value: isinstance(x, list) / isinstance(x, dict) still hold, so unification and clause-head indexing see no difference from an ordinary literal; only mutation is blocked. This protection is not limited to the module-global-held original: every RECONSTRUCTION a compiled clause builds when it references the constant is frozen too, the same way — load-bearing under tabling, where a cached answer can be shared across multiple consumers, and one consumer mutating it would corrupt what every other consumer of that same cached answer sees. Constants cannot be hidden from Python at module level — the module global holds this same frozen value, and module_constant/3 reflects it (see Builtins) — frozen, not hidden. Need a mutable working copy instead? copy.deepcopy(constant) (or list(constant) / dict(constant) / set(constant) for a shallow one) returns a PLAIN, unfrozen container — that round-trip (also how pickle serializes a frozen constant) is the documented escape route. The one deliberate exception to all of this: a functor constant's own field values are not frozen (point([1, 2, 3], 0)'s list field can still be mutated) — freezing stops at the term boundary a functor call introduces, not inside it. And one type-changing freeze: a ++()-escape-built bytearray value freezes to bytes (a different type, not a frozen subclass) — bytes already is Python's immutable byte-string type, so there is nothing to subclass.

A bare reference folds; ++name looks up

A bare constant reference is replaced by its ground value during compilation — the same mechanism that already folds the true/false/undefined truth-value aliases. There is no Var, no deref, and no runtime cost: pi in a compiled clause is 3.14159. This applies uniformly, including in head position — area(pi, R) compiles exactly as area(3.14159, R) would, and dispatching on the literal value is a legitimate idiom that earns no lint.

++pi does the opposite and is usually what you want: it resolves the module global when the goal runs, so a constant is late-bound and one declaration governs every use. Rebinding the global after load moves a ++pi answer and leaves a folded pi answer where it was.

Referencing a name that nothing declares is the ordinary strict-atoms NameError, the same one any other undeclared bare name gets. There is no constants-specific diagnostic any more: the name no longer says it is a constant, so nothing can tell a mistyped max_fien from an ordinary Python helper. Inside a ++() or f-string escape the body is verbatim Python, so a free name there fails as a Python NameError — at load time in a -constant_value RHS, which is evaluated at module level, and when the goal runs in a clause body.

Importing

Constants export automatically — every -constant_value declaration is a public module global, so there is nothing to list in -module or -private. Listing the name there is not an error, though: it declares the ATOM of the same spelling, which is the intended pairing — pi the atom and ++pi the value, in one file. Import with the same directives used for predicates:

-import_from(other_module, [pi])                          # direct
-import_from(other_module, [alias(pi, mypi)])              # alias — an ordinary rename
-import_module(other_module)                               # whole-module, qualified access
circumference(R, C) <- (C == 2 * other_module.pi * R)      # bare-term — works directly

Qualified access (other_module.pi) works both as a bare term and inside ++() — the same dotted-attribute mechanism that already resolves a qualified atom reference like currency.euro covers constants too. See Import System for the full directive semantics.

The _UNUSED edge

A constant name ending in _UNUSED (item_UNUSED) is legal, but earns a load-time ClausalLintWarning — it visually collides with the singleton-suppression suffix above, which applies to variables, not constants. Pick a different name.


Atoms

Inside a logical term: - lowercase identifiers that are not logic variable names are atoms

(TitleCase was an atom spelling too until 2026-09-10. It is a logic variable now, so a bare atom is written lowercase, and a bare TitleCase name in functor position is a load-time error rather than an atom — see the asymmetry noted there. Single-quoting is unaffected: 'Foo' is the atom Foo in every position, functor position included.)

# These are atoms (bare identifiers; unicode letters OK). Each is the
# interned str 'red' / 'active' at runtime:
color(red),
status(active),
# Single quotes spell an atom whose name is not a bare identifier:
grade('A+'),
# A double-quoted literal is an atom or a string per the file's
# -double_quotes mode (see Strings):
greeting("hello world"),

An identifier that collides with a Python keyword or builtin (not, is, max) cannot be written as a bare atom. Single-quote it — 'not' — and you get exactly the atom not, in every -double_quotes mode (see Atoms vs strings). Single quotes are also how you spell an atom whose name has spaces or punctuation: 'hello world', 'order #42'.

An atom is the interned Python str — no wrapper, no class: red and 'red' both produce the same str object every reference to red produces. Atoms are first-class values you can pass around, store in dicts, and use as goals. They are compared with == (str equality); because atoms are interned, is happens to agree too, but == is the test to write. From Python, build and read them with clausal.logic.atoms:

from clausal.logic.atoms import mint, is_atom, spelling

mint("red")            # → 'red'
is_atom("red")         # → True
spelling("red")        # → 'red'

The 1-tuple ("red",) — how atoms used to be represented — is now reserved for a future opaque Python object reference and is refused with a TypeError if constructed as a term.

Strict by default; global identity when resolved

An undeclared bare atom reference is a compile-time NameError by default — matching Python's treatment of undefined names. A bare red that has not been declared (via -module, -private, or -import_from) or reached via global_atom/2 will not compile.

There is no per-file opt-out: -implicit_atoms was removed before 1.0 — declare the atoms, or quote them.

Once an atom is resolved — declared in -module([...]) (public), -private([...]) (private), imported, or reached via global_atom/2 — it is global by spelling: every module that writes red has the same atom, the interned str 'red', and the two compare equal, matching Prolog's convention. == is the test to write.

# module_a.clausal — bare reference, no -module/-private listing
color_a(red),

# module_b.clausal — bare reference, no -module/-private listing
color_b(red),

# Both files' `red` are the same atom, global by spelling.
# In Python:
#     from module_a import red as red_a
#     from module_b import red as red_b
#     assert red_a == red_b        # value equality — never `is`
#     assert red_a == mint("red")  # clausal.logic.atoms.mint

-module([...]) and -private([...]) differ only in whether the name is advertised as part of the module's surface — a -private atom is still importable, and neither changes the atom's spelling. For an atom another module genuinely cannot reach, use -hide([...]), which renames it to a spelling the reader refuses to accept:

# -module and -private only say whether the name is part of the module's
# advertised surface; the atom is `red` either way, global by spelling.
-module(traffic, [red, amber, green])

# OR: internal to this module, still an ordinary global `red`.
-private([red])

# For an atom no other module can reach or spell, use -hide (requires a
# preceding -module): the compiler renames it into the module's namespace.
-hide([sentinel])

Resolution order for a bare reference inside a module: -private → -module → -import_from → global fallthrough.

The full design is in the global-atoms-default spec.

Predicates with arity ≥ 1 do not participate in the global default: they remain module-local-by-default and must be listed in -module / -private (or imported via -import_from) to be shared across modules. The asymmetry between atoms (global default) and predicates (local default) is principled — Prolog itself treats atoms as global and predicates as module-scoped.


Builtin predicate naming

Built-in predicates are lowercase snake_case (e.g. findall, assertz, var, read_file). Where a natural name collides with Python, a trailing underscore disambiguates. This is a deliberate design choice — not aesthetic — with two goals:

  1. Avoid Python keyword conflicts. Many natural predicate names are Python reserved words: in, is, not, and, or, if, for, assert, lambda, global, return, yield. A predicate named in would be a syntax error the moment it appears as a call in a clause body — so membership is spelled in_.

  2. Avoid Python builtin conflicts. Names like abs, all, any, filter, float, int, map, max, min, set, str, sum are Python builtins that would shadow (or be shadowed by) a same-named predicate — so these take a trailing underscore: abs_, float_, max_, min_, sum_, divmod_.

The trailing underscore keeps the builtin namespace cleanly separate from Python keywords and builtins. User predicates follow the same snake_case convention, and only need the underscore where they would hit the same collision.

  1. Engine-provided type names are injected. For convenience in embedded Python, Var, Trail, DictTerm, SetTerm, PyThunk, Quantity and Undefined are injected into every module's namespace. They are TitleCase, so they never collide with a predicate name (a TitleCase functor is a load-time error). The engine functions walk, deref, and unify are injected under an internal $-prefix, so walk/2, deref/2, and unify/2 are free for user predicates.

Unification

Unification is written with is:

X is 42,
X is Y,
X is not Y,   # disequality constraint (dif/2)

A bare is is unification, never evaluation. X is 1 + 2 binds X to the term 1 + 2, not to 3. Prolog's arithmetic is/2 is written quoted, 'is'(X, 1 + 2), or as eval_(1 + 2, X); both bind X to 3. The quoted spelling is the ISO predicate, as every quoted operator is: see Operators for the full bare-versus-quoted table, and Arithmetic for evaluation.

test("bare is unifies") <- (X is 1 + 2, not (X is 3))
test("quoted is evaluates") <- ('is'(X, 1 + 2), X == 3)
test("quoted = unifies too") <- ('='(X, 'a'), X is 'a')

Why is rather than =? - = is Python's assignment operator and cannot appear in a clause body (the quoted '='(A, B) is ISO unification) - is expresses the same concept in English — two things being the same — and Python programmers understand it. The seam generalises this concept; in 'X is Y', if we don't know X or Y, we are describing that they must be same whatever they are, and when either become known, they both become known.

Conversely, X is not Y posts a disequality constraint (dif/2): X and Y must end up with different values. This is lazily checked — the constraint is re-evaluated each time either variable gets bound. If they become equal, the constraint fails and the search backtracks. If they remain different, the constraint is satisfied and dropped. See constraints.md for details.

not (X is Y) is the immediate check (ISO '\\='(X, Y)): it fails if X and Y can unify right now, regardless of future bindings. Use this when you want point-in-time semantics.

The corresponding AST node is Unify(left, right). Disequality is DoesNotUnify(left, right).

Inline naming with is-chains

Python's comparison chaining gives the seam unification chains: A is B is C unifies pairwise, and the shared middle operand is evaluated once. This names a term and uses it in the same goal — where Python code would reach for the walrus operator:

-allow_singletons
# VALUE is named to demonstrate the inline-naming feature itself — that
# it *can* be named is the point, not any further use of it here.
test("name a term inline") <- (
    D is {'k': [1, 2]},
    VALUE is [1, X] is D['k'],
    X == 2
)

Here [1, X] is constructed once, named VALUE, and unified with D["k"] — binding X to 2 in the process. Chains of any length work; each link is an independent unification of the adjacent operands.


Arithmetic binding

To constrain a variable to an arithmetic expression, use ==:

fib(N, RESULT) <- (
    N > 1,
    N1 == N - 1,
    N2 == N - 2,
    fib(N1, A),
    fib(N2, B),
    RESULT == A + B
)

N1 == N - 1 posts an arithmetic constraint relating N1 and N. Unlike Prolog's is/2, this works in all directions — even when N is unbound. No extra parentheses are needed: clause bodies are already inside (...). This allows sophisticated reasoning about numbers, using the builtin constraint logic programmming (CLP) modules.

The distinction from is: - X is Y — pure structural unification; neither side is evaluated arithmetically - (X == expr) — posts an arithmetic constraint (CLP(ℤ) or CLP(ℝ)) - eval_(expr, X), or the quoted ISO form 'is'(X, expr) — eager arithmetic evaluation; evaluates expr and binds the result to X (Prolog's is/2). A variable operand is evaluated at runtime, whatever arithmetic term it holds; an unbound one raises instantiation_error, and an atom or a non-evaluable compound raises type_error(evaluable, Name/Arity). See Arithmetic for the evaluable table and for when to prefer it over ==.


Comparison operators (CLP(ℤ))

The comparison operators ==, !=, <, >, <=, >= are CLP(ℤ) (Constraint Logic Programming over Integers) operators. They post constraints on integer variables rather than performing immediate checks.

bounded(X) <- (
    in_domain(X, 1, 10),
    X > 3,
    X < 8,
    label([X])
)

when both sides are ground (no unbound Vars), the operators fall back to direct Python comparison — 3 == 3 is True, 3 < 2 is False — so existing ground arithmetic code works unchanged.

when at least one side is an unbound Var, a CLP(ℤ) constraint is posted: - X == 5 narrows X's domain to {5} (and binds it) - 1 <= X and X <= 10 constrain X's domain to [1, 10] - X != 3 removes 3 from X's domain - X < Y narrows X's upper bound and Y's lower bound

Operator CLP(ℤ) meaning
== Arithmetic equality constraint
!= Arithmetic disequality constraint
< > <= >= Comparison constraints (narrow domain bounds)

== and != post arithmetic constraints (clpz's #=/2 and #\\=/2); inside a constraint / is exact rational division. For structural identity (ISO ==/2) — comparing terms without binding or evaluating — write the quoted '=='(X, Y) (or structural_eq(X, Y)), and '\\=='(X, Y) for its negation. The quoted ISO arithmetic comparisons ('=:=', '<', '=<', ...) evaluate both sides and post nothing. Operators has the full table.

See constraints.md for the full CLP(ℤ) design, including domain representation, propagation, and labeling.

Rational constraints: a set in goal position

A set literal in goal position is a set of CLP(ℚ) constraints over exact rational arithmetic — the twin of Prolog's clpq goal {C}. Its elements are comparisons (==, !=, <, <=, >, >=, or a chain), posted together:

two(N, Q) <- {Q == N * 2}                    # two(3, Q) gives 6; two(N, 8) gives 4
within(X) <- {0 <= X <= 10}
half(X) <- {2 * X == 3}                       # X = Fraction(3, 2)

It is one goal, clpq.rational((...)), spelled short. A set in data position (a head, an argument, S is {1, 2}) is still a set, and {} is a dict. An element that is not a comparison is a load-time error. See CLP(Q).


Horn clauses

head(X) <- (goal1(X), goal2(X))

The <- operator denotes a Horn clause (rule). It will never be added to Python's expression grammar because it conflicts with x < -y (less-than applied to a negated value) — but only when there is no surrounding whitespace. With whitespace, it is unambiguous and parseable.

Body style

The body after <- must be one of:

  • A single call — no parentheses needed:

    sorted_asc([_]),
    palindrome(XS) <- reverse(XS, XS)
    

  • A bare name — no parentheses needed:

    always_true <- True,
    

  • Anything else — parenthesized:

    safe_max(X, Y, X) <- (X >= Y)
    fib(N, RESULT) <- (
        N > 1,
        N1 == N - 1,
        N2 == N - 2,
        fib(N1, A),
        fib(N2, B),
        RESULT == A + B
    )
    

This rule exists because Python's parser sees <- as < followed by unary -. when the body contains operators (+, <, and, or, not, etc.), the - gets absorbed into the body expression and the AST is silently mangled. Parentheses force Python to treat the body as a single grouped expression, keeping the - at the top where the term rewriter can find it. Calls and bare names are safe without parentheses because they bind tighter than unary -.

To keep things safe, attempting to write an unparenthesized operator body produces a clear error:

SyntaxError: clause body must be parenthesized or a single call:
    write  head <- (body)  or  head <- goal(X)

Conjunction style

Multiple goals in a body are separated by commas, with each goal on its own line:

is_permutation(XS, YS) <- (
    length(XS, N),
    length(YS, N),
    sort(XS, S),
    sort(YS, S)
)

This looks a bit like a Python function def doesn't it? But it is actually much more powerful. These clauses describe a relation. This can go in multiple directions, and this generality is fundamental to the power of logic programming.

As usual in programming, be careful about operator precedence: Inside sub-expressions like not (...) or ... or ..., use and instead of commas — commas inside these would be parsed as Python tuples:

test("fails") <- (not (X is 1 and X is 2))
test("either") <- (X is 1 or X is 2)

What if we just want to state a fact that always holds? We could do so by using a body that is always true, i.e. True.

parent('tom', 'bob') <- True

But there is a shorthand for this. Facts (trivially true rules) are simply written without a body, only a trailing comma:

parent('tom', 'bob'),
parent('bob', 'ann'),

(The atoms are quoted because a bare tom must be declared first, e.g. -private([tom, bob, ann]); see Atoms.)

Grammar rules (Definite Clause Grammars):

rule >> list_description

Lists

[]               # empty list (singleton)
[a, 1, X]        # a simple list
[FIRST, *REST]   # head/tail decomposition
[*BEFORE, PIVOT, *AFTER]  # multiple spread patterns

Partial lists (Prolog [H|T] where T is a variable) use Python's * spread syntax rather than |. The empty list is a singleton — unlike Python, two [] literals are the same object.


Dicts

Python dict literals in seam (.seam) files create DictTerm objects — unification-aware dictionaries. Keys must be ground; values may be logic variables.

# Ground dict fact
point({'x': 0, 'y': 0}),

# Dict pattern in head — X binds during unification
get_x({'x': X, 'y': _}, X),

# Dict construction in body
make_point(X, Y, P) <- (P is {'x': X, 'y': Y})

# Nested dicts
get_city({'address': {'city': C}}, C),

Two dicts unify iff they have the same keys and values unify pairwise. A variable unifies with a dict by binding to it.

See Dicts & Sets for the full design.


Sets

Python set literals in seam (.seam) files create SetTerm objects — unification-aware sets. Elements must be ground (hashable).

colors({1, 2, 3}),
primary({'red', 'green', 'blue'}),

Two sets unify iff they contain the same elements (order irrelevant). Variables in set elements are not supported.

This is the data position. A set literal standing alone as a goal is not a value but a CLP(ℚ) constraint set: {X >= 0, X <= 10}.

See Dicts & Sets for details.


Strings

"…" is a string by default, as in Scryer and Trealla: the list of its one-character atoms — the classical Prolog chars model (ISO's double_quotes flag set to chars). A bare Python str is an atom, not a string — see Atoms vs strings — and a string reaches Python as the carrier ('$chars', text). -double_quotes(atom) is a temporary per-module setting for code not yet migrated. See strings as lists:

"abc" is ['a', 'b', 'c'],  # a string IS its char list
"" is [],                  # the empty string is the empty list
"abc" is [H, *T],          # H = 'a' (a char atom), T = "bc" (a string)

All list operations apply to strings, because a string is a list: append/3, length/2, reverse/2, list_item/3, in_/2, maplist/N, DCGs. "" and [] are one and the same term. A string is not an atom — atom/1 rejects it, string/1 and is_str/1 accept it, and is_list/1 accepts it too. The analogous b"…" byte literal is a list of integer codes (0–255); str and bytes are distinct domains and never cross-unify.

Under the hood a string stays a compact value — it is never expanded into a chain of cons cells — but every relation treats it as the char list it denotes.

Which quote you write decides what you get

Unlike every earlier release, the quote character is now significant: 'foo' is always the atom foo; "foo" is a string (unless the file sets -double_quotes(atom)). The u, r and triple-quote prefixes are inert — the quote character after any prefix is what counts. b"…"/b'…' are always codes, in either quote style.


Atoms vs strings

Clausal Prolog keeps two disjoint kinds, exactly as ISO Prolog does:

Atom (a symbol) String (text / data)
written as bare identifier red, café, δικαίωμα; or 'any spelling' "hello world" (the default)
represented as the interned Python str itself the carrier ('$chars', text), which every relation treats as its char list (a '.'/2 compound) — never a bare str
compared by == (str equality; interning makes is agree too, but write ==) value equality; also unifies with its char list
typo-safe? yes, under -strict_atoms (the default) no (it's data)
atom/1 matches does not match (use string/1 / is_str/1)
atomic/1 matches does not match — a string is a list
callable_/1 matches (an atom can name a goal) matches when non-empty — it is a '.'/2 compound, as in ISO ("" is [], an atom)
indexed on? yes — first-argument indexing keys ('red', 0) no — a string head argument falls in the full-scan bucket

The two never unify: red = "red" fails. A one-character string is not a character either — "a" is the list [a], while a is the char atom, so "a" = a fails as well.

Quoting is how you spell an atom whose name is not a bare identifier. 'hello world', 'order #42' and 'not' are ordinary atoms — the single quote works in every -double_quotes mode, including for names that collide with a Python keyword or builtin.

A string is never a functor. "foo"(1) is a SyntaxError in every mode (the ISO functor rule); write 'foo'(1) — or a bare foo(1) — when you mean the compound term.

Choosing between them

Reach for an atom when the value is a symbol the program reasons about: a tag, a status, a colour, a key. It is typo-safe under strict atoms, it is cheap to compare, and it participates in first-argument indexing. Reach for a string when the value is text that came from, or is going to, the outside world: a file line, a JSON value, a message. Text that crosses to Python and back comes back as a string (Python integration), so this is also the direction the boundary pushes you.

If you need an atom built at runtime — interning symbols imported from an external system, say — use atom_chars/2 on the text, or Python's clausal.logic.atoms.mint("any name").


F-strings

Python f-strings work naturally in seam (.seam) files. Logic variables are auto-dereferenced at search time — bound variables interpolate their value, unbound variables show _N.

greet(NAME) <- writeln_text(f"Hello, {NAME}!")

show_pair(X, Y) <- writeln_text(f"{X} and {Y}")

# Format specs work too
show_price(ITEM, PRICE) <- writeln_text(f"{ITEM}: ${PRICE:.2f}")

Under the hood, f-strings in .seam files are compiled to deferred PyThunk lambdas during AST transformation. Logic variable names become lambda parameters; the compiler emits calls with deref()'d values at search time.

Simple variable references like f"{X}" and f"{NAME}" work correctly. Format specs (:.2f, :>10, etc.) and conversions (!r, !s) are fully supported. Python expressions inside f-strings (like f"{len(L)}" or f"{S.upper()}") also work — the entire f-string is wrapped in a lambda that receives dereferenced values.


Marking a variable inside a thunk

An f-string slot and a ++ operand are verbatim Python, so a name written there could mean either the clause's logic variable or a binding in the module namespace. Written bare, the reading is decided for you: a name the clause uses as a variable elsewhere is captured, and any other name resolves in the module namespace when the thunk runs.

--X states it instead. Inside an f-string slot or a ++ operand, --X means the logic variable X:

label(S) <- (tree(Node), S is f"{--Node}")
shout(S) <- (tree(Node), S is ++str(--Node).upper())

This is purely additive — bare X keeps working exactly as before, and both spellings give the same answer wherever the marker is accepted. What to know:

  • It is recognised anywhere inside the slot or operand, not only at the top: a thunk body is usually a call, and ++len(--List) is the shape that matters. As everywhere else, the two - must be adjacent — - -X is double negation, which is also how a <- -b stays an arrow.
  • It is checked. If no goal outside a thunk uses the name, --X is a load-time error naming both readings. A bare X cannot be checked this way, because a bare name the clause does not bind is a legitimate reference to the module namespace; a marked one is not.
  • Only an identifier that is a variable spelling is a marker. --total is the double negation it always was.
  • It is not recognised inside an inline -- seam. There a ++ operand is hosted Python again and --expr is already the seam itself, nesting to any depth, so the marker would be a second meaning for one spelling. The with --{} block form is not a seam operand — its statements are logic terms — and markers do work there.
  • Not in a format spec. A format spec is a STRING, not a term position — it describes how to render a value, it does not name one. So neither spelling captures a clause variable there, and that is deliberate rather than a gap: f"{N:>{W}}" resolves W in the module namespace, exactly as the surrounding Python would.

The marker, however, promises that it cannot be silently misread, so f"{N:>{--W}}" is a load error rather than a marker that quietly does nothing. Put the marker in a value slot — f"{--Node:>8}" is fine, the restriction is on the spec, not on formatting — or do the whole formatting inside a ++ escape.

The asymmetry is the point: the bare spelling behaves as Python does, and the marker refuses rather than pretend.


Python interop — ++() escape

The ++() operator evaluates an arbitrary Python expression at search time. Logic variables inside the expression are automatically dereferenced.

As a value (inside is):

# Call a Python builtin
list_len(L, N) <- (N is ++len(L))

# Method call on a dereferenced variable
to_upper(S, R) <- (R is ++S.upper())

# Arithmetic
inc(X, R) <- (R is ++(X + 1))

# Subscript access
first(L, R) <- (R is ++L[0])

# Dict access
get_key(D, K, R) <- (R is ++D[K])

# Multiple logic variables
add_len(A, B, R) <- (R is ++(len(A) + len(B)))

# No logic variables (pure Python)
get_pi(R) <- (R is ++(3.14159))

As a goal (side effects):

# Print as a goal
show(X) <- ++print(X)

# Goal followed by continuation
process(X, R) <- (
    ++print(X),
    R is ++(X * 2)
)

Under the hood, ++expr wraps the Python expression in a lambda whose parameters shadow the module-scope Var names. The compiler emits thunk_fn(deref(v0), deref(v1), ...). Any Python expression works — method calls, builtins, arithmetic, subscripts, etc.

The values cross in both directions without a converter: a string argument ("abc") reaches the Python code as a plain str, and a str the expression returns is an atom — to_upper("abc", R) binds R to the atom 'ABC'. Anything else passes through unconverted (see Python integration).

Unit-literal sugar — n(Unit)

A special case of the ++() pattern: when a numeric literal is used as the callable with a single unit-predicate argument, it desugars to ++(Unit(n)):

5(metre)          # → Quantity(5, metre)
9.8(newton)       # → Quantity(9.8, newton)  (dimension kg·m·s⁻²)
-3(second)        # → Quantity(-3, second)  (negation applied after)

when a logic variable is used as the callable instead, X(Unit) becomes a goal that posts a dimension constraint on X:

FORCE(newton)         # → has_units(FORCE, newton) — FORCE must be a newton value
FORCE is 9.8(newton)  # binds FORCE; the hook checks the dimensions match

See Units for the full reference.


Compound terms and goals

goal(_, _),             # compound goal
not goal,               # negation as failure

A compound term is a cell: the plain tuple (functor, *args), so point(1, 2) is ('point', 1, 2) to Python code (read it with clausal.cell_functor / clausal.cell_args). Arguments are positional. A keyword-argument term is a load-time SyntaxError:

-private([point(x, y)])
p(P) <- (P is point(x=1, y=2))

SyntaxError: ... `point/2` is written with keyword arguments (x=, y=): a term is
built positionally. Write the arguments in the declared order, ...

The field names in a declaration (point(x, y)) are still read by vary/3, unbound_keys/2 and signature/3. Two keyword spellings remain: a directive's options (-specialize(solve, p, alias=q)) and an EDCG hidden argument (_edcg_len_in=0).


Immediate goals (planned)

Note: This syntax is not yet implemented. Use the assertz(goal) and retract(term) builtins directly.

+ goal,     # assert/call immediately ('+' distinguishes from a fact)
- term,     # retract term

Module qualification

Predicates from imported modules are called with dotted notation after loading the module:

utils.double(X, Y),    # qualified call after -import_module(utils)

See Directives and Import System for details.


Lambdas

Lambdas are anonymous clauses — goal closures passed as arguments to higher-order predicates. They use the same head <- body arrow syntax as clause definitions:

# One-arg lambda — X is a parameter, RESULT is captured
apply(RESULT, VAL) <- call_goal((X <- (RESULT == X + 1)), VAL)

# Two-arg lambda
apply_add(A, B, R) <- call_goal(((X, Y) <- (R == X + Y)), A, B)

# Zero-arg lambda
run_goal(RESULT) <- call_goal((() <- (RESULT is 42)))

# Captured variable from enclosing clause
add_z(Z, R) <- call_goal((X <- (R == X + Z)), 10)

# Conjunction body
transform(R) <- call_goal(((X, Y) <- (T == X + 1, Y == T * 2)), 5, R)

Parameters are lambda arguments; captured variables share the enclosing clause's Var objects. Body-local variables (first appearing inside the lambda) get fresh Var() allocations. Lambdas are called via the call_goal/1..8 builtins (or call/1..8).

See Lambdas for the full design, compilation details, and examples.


Definite Clause Grammars — >>

DCG rules provide syntactic sugar for difference-list grammars. Each >> rule compiles to an ordinary <- clause with two extra hidden arguments (input list, remaining list) threaded through the body. This is the same approach as Prolog's -->, using Python's >> operator instead.

Basic syntax

# Terminal — consume literal tokens from the input list
greeting >> (['hello', 'world'])

# Non-terminal — call another DCG rule (state threaded automatically)
sentence >> (noun_phrase, verb_phrase, noun_phrase)

# Empty terminal (epsilon — matches without consuming)
epsilon >> ([])

Extra arguments and inline goals

DCG predicates can have extra arguments beyond the hidden state:

# Extra arg D, plus inline goals {D >= 0} and {D <= 9}
digit(D) >> ([D], {D >= 0}, {D <= 9})

Inline goals are written with {...} (Python set literal syntax). They execute without consuming input — the state passes through unchanged. Multiple consecutive inline goals are optimised to avoid generating unnecessary intermediate state variables.

Conjunction and disjunction

# Conjunction — comma-separated (canonical style)
rule >> (a, b, c)

# Disjunction
letter >> (['a'] or ['b'] or ['c'])

Negation

# Negation as failure — state passes through
not_a >> (not ['a'], [X])

Pushback / semicontext

The LHS can be a tuple (head, [pushback_tokens]) to push tokens back onto the input after matching:

# Peek at next token without consuming it
(look_ahead(T), [T]) >> ([T])

After the body matches [T], the pushback [T] is prepended to the remainder.

Invoking DCGs with phrase

Use phrase/2 or phrase/3 to call DCG rules from regular predicates:

# phrase/2 — must consume the entire input list
valid_sentence(S) <- phrase(sentence, S)

# phrase/3 — partial parse, remaining input bound to REST
phrase(digit(D), [3, 'plus', 4], REST)

phrase/2 passes [] as the expected remainder, so the rule must consume all input. phrase/3 leaves the remainder as a logic variable for partial parsing.

Module exports

when using -module(...), DCG predicates must be declared with their full signature including the two hidden state arguments:

# Correct: predicates declared with proper arities
-module(my_grammar, [greeting(S0, S), digit(D, S0, S)])

# Wrong: greeting/digit are read as atoms here, not predicate references
-module(my_grammar, [greeting, digit])

How it works

The >> rewriting is purely syntactic — it transforms DCG rules into ordinary <- clauses before the compiler sees them:

# This DCG rule:
greeting >> (['hello', 'world'])

# Rewrites to this ordinary clause:
greeting(S0, S) <- (S0 is ['hello', 'world', *S])
# This DCG rule:
digit(D) >> ([D], {D >= 0}, {D <= 9})

# Rewrites to:
digit(D, S0, S) <- (S0 is [D, *S], D >= 0, D <= 9)

No changes to the compiler, database, or runtime are needed.

DCGs as general state-passing

The hidden list pair can thread any state, not just tokens: pass [Initial] to phrase/3 and receive [Final], with two helper nonterminals reading and replacing the state. DCGs — state threading has the pattern and worked examples (counter, tree leaves, accumulator):

(state(S), [S]) >> ([S])            # read the state
(state2(S0, S), [S]) >> ([S0])      # read the old state, write a new one
inc >> (state(N0), {N == N0 + 1}, state2(_, N))

count3 >> (inc, inc, inc)

test("count from 0") <- phrase(count3, [0], [3])

(Give the two helpers different names: state/1 and state/2 as pushback nonterminals in one module currently fail to load.)


Extended DCGs — EDCGs

Standard DCGs thread a single state (the difference list). Extended DCGs add support for multiple named accumulators and read-only passed arguments, all threaded automatically through >> rules. This eliminates the boilerplate of manually encoding multiple states into a single compound value.

EDCGs are based on Peter Van Roy's 1989 design and use three directives to declare the threading:

Declaring accumulators

An accumulator has a name and a joiner goal that relates a pushed value to the input/output state:

# Numeric counter: Out = in + Value
-edcg_acc(counter, X, IN, OUT, {OUT == IN + X})

# List accumulator: prepend items
-edcg_acc(items, ITEM, IN, OUT, {OUT is [ITEM, *IN]})

# Product accumulator: Out = in * Value
-edcg_acc(product, X, IN, OUT, {OUT == IN * X})

The joiner goal can be any clausal goal wrapped in {braces}. The variable names (X, IN, OUT) are placeholders — they get substituted with actual variables during rewriting.

Declaring passed arguments

A passed argument is a read-only value threaded unchanged through all sub-calls:

-edcg_pass(config)
-edcg_pass(scale)

Declaring predicates

Each EDCG predicate must declare its visible arity and which accumulators/passes it uses:

-edcg_pred(inc, 0, [counter])           # 0 visible args, uses counter
-edcg_pred(process, 1, [counter, items]) # 1 visible arg, uses counter + items
-edcg_pred(parse, 0, [counter, dcg])     # uses counter + standard DCG list
-edcg_pred(scaled_inc, 0, [counter, scale])  # accumulator + passed arg

The special name dcg refers to the standard DCG difference-list accumulator. Include it when your EDCG rule also parses tokens.

EDCG rule syntax

EDCG rules use >> just like standard DCGs, with additional operators:

# Push a value to a named accumulator: [value] // acc_name
inc >> ([1] // counter)

# Read current accumulator value: acc_name / Var
get_and_inc(V) >> (counter / V, [1] // counter)

# Read a passed argument: pass_name / Var
scaled_inc >> (scale / S, [S] // counter)

# Terminal list (requires 'dcg' in the predicate's accumulator list)
token(T) >> ([T], [1] // counter)

# Inline goals don't thread accumulators
inc_if_positive >> (counter / N, {N >= 0}, [1] // counter)

# Sub-calls: accumulators are threaded automatically
count3 >> (inc, inc, inc)

# Empty body: all accumulators pass through unchanged
noop >> ([])

The // operator pushes a value through the accumulator's joiner goal. The / operator reads the current state without modifying it.

Multiple accumulators

A single rule can update multiple accumulators simultaneously:

-edcg_acc(counter, X, IN, OUT, {OUT == IN + X})
-edcg_acc(items, ITEM, IN, OUT, {OUT is [ITEM, *IN]})
-edcg_pred(process, 1, [counter, items])

# Each push targets a specific accumulator by name
process(X) >> ([1] // counter, [X] // items)

when a sub-call uses fewer accumulators than the caller, only the shared ones are threaded:

-edcg_pred(inc_only, 0, [counter])          # only counter
-edcg_pred(do_both, 1, [counter, items])    # counter + items

inc_only >> ([1] // counter)
do_both(X) >> (inc_only, [X] // items)    # inc_only threads counter only

Calling EDCG predicates

EDCG predicates are compiled to ordinary predicates with hidden arguments appended in declaration order: 2 per accumulator (in, out) + 1 per pass. You can call them from regular <- clauses using keyword syntax:

# -edcg_pred(count_elems, 1, [len])
# Compiled arity: 1 (visible) + 2 (len_in, len_out) = 3
my_length(L, N) <- count_elems(L, _edcg_len_in=0, _edcg_len_out=N)

Or positionally — hidden args follow visible args in the order declared:

# count_elems(List, len_in, len_out)
my_length(L, N) <- count_elems(L, 0, N)

Complete example: counter with scale factor

-module(example, [run_scaled(LIST, SCALE, COUNT, ITEMS)])

-edcg_acc(counter, X, IN, OUT, {OUT == IN + X})
-edcg_acc(items, ITEM, IN, OUT, {OUT is [ITEM, *IN]})
-edcg_pass(scale)

-edcg_pred(scaled_inc, 0, [counter, scale])
-edcg_pred(collect_and_count, 1, [counter, items, scale])
-edcg_pred(process_list, 1, [counter, items, scale])

scaled_inc >> (scale / S, [S] // counter)
collect_and_count(X) >> (scaled_inc, [X] // items)

process_list([]) >> ([])
process_list([X, *XS]) >> (collect_and_count(X), process_list(XS))

run_scaled(LIST, SCALE, COUNT, ITEMS) <- (
    process_list(LIST, _edcg_counter_in=0, _edcg_counter_out=COUNT,
                 _edcg_items_in=[], _edcg_items_out=ITEMS,
                 _edcg_scale=SCALE)
)

Control flow in an EDCG body

Disjunction (or), negation (not) and reified if-then-else (if_/3) all thread accumulators:

-module(edcg_ite, [classify(_edcg_counter_in, _edcg_counter_out)])
-edcg_acc(counter, X, IN, OUT, {OUT == IN + X})
-edcg_pred(inc, 0, [counter])
-edcg_pred(classify, 0, [counter])

inc >> ([1] // counter)
classify >> (if_({1 == 1}, inc, (inc, inc)), inc)
  • Disjunction and if_/3 are joins: every branch is rewritten from the same starting state and meets at one variable, so a branch that pushes fewer times than its siblings is padded out. Whatever follows the construct continues from that meeting point.
  • The if_/3 condition is a {Goal} block: it touches no accumulator, and Goal must be reifiable (a comparison or a reified closure; see Reified if-then-else). A condition that pushes, or a non-terminal as the condition, is refused at load time: it would be a plain goal, which if_/3 does not take (ruling 2026-10-01).
  • Negation does not consume: not G leaves every accumulator where it found it.

Design notes

  • Purely syntactic: EDCG >> rules are rewritten to ordinary <- clauses before compilation. No runtime support needed.
  • Backward compatible: Rules without -edcg_pred declarations continue to use standard DCG rewriting.
  • // for push, / for read: These use Python's floor-division and division operators respectively.
  • Accumulator order matters: Hidden args are appended in the order listed in -edcg_pred. when calling from <- clauses, match this order.

Meta-predicates

Meta-predicates are higher-order predicates that take goals as arguments. They are compiled as special forms — the goal argument is compiled inline, not passed as a runtime value.

All-solutions predicates

test("findall") <- findall(X, in_(X, [1, 2, 3]), [1, 2, 3])
test("filter") <- findall(X, (in_(X, [1, 2, 3]) and X > 1), [2, 3])
test("product") <- findall([X, Y], (in_(X, ['a', 'b']) and in_(Y, [1, 2])), [['a', 1], ['a', 2], ['b', 1], ['b', 2]])
test("bagof empty fails") <- (not bagof(X, in_(X, []), _))
test("findall empty") <- findall(X, in_(X, []), [])
test("setof sorts and dedups") <- setof(X, in_(X, [3, 1, 2, 1]), [1, 2, 3])
Predicate Empty result
findall/3 Succeeds with Bag = []
bagof/3 Fails
setof/3 Fails

Universal quantification

# forall(Cond, Action) succeeds iff Action holds for every solution of Cond
test("forall holds") <- forall(in_(X, [2, 4, 6]), X > 0)
test("forall fails") <- (not forall(in_(X, [2, -1, 6]), X > 0))

forall(Cond, Action) is equivalent to not (Cond and not Action).

call/N

call/N invokes a goal closure with extra arguments. It is an alias for call_goal/N:

call_goal((X <- (X > 0)), 5),        # call_goal/2: succeeds
call(GOAL, ARG1, ARG2),             # call/3: invoke GOAL with two extra args

call/1 through call/8 are available (as are call_goal/1 through call_goal/8).

Higher-order list predicates

These predicates take a goal closure and apply it across a list. All use committed choice (first solution per element).

# maplist/2 — check Goal(Elem) succeeds for every element
maplist((X <- (X > 0)), [1, 2, 3]),              # succeeds

# maplist/3 — map Goal(X, Y) over list, collect results
maplist(((X, Y) <- (Y == X * 2)), [1, 2, 3], YS),  # YS = [2, 4, 6]

# include/3 — keep elements where Goal(Elem) succeeds
include((X <- (X > 0)), [1, -2, 3, -4], R),      # R = [1, 3]

# exclude/3 — keep elements where Goal(Elem) fails
exclude((X <- (X > 0)), [1, -2, 3, -4], R),     # R = [-2, -4]

# foldl/4 — left fold with Goal(Elem, Acc0, Acc1)
foldl(((E, A, R) <- (R == A + E)), [1, 2, 3], 0, SUM),  # SUM = 6

Constraint logic programming

Clausal Prolog supports CLP(ℤ) (integer constraints) and CLP(B) (Boolean constraints). Constraint operators are used directly in clause bodies — no special escape or domain wrapper is needed.

in_domain(X, 1, 9),
all_different([X, Y, Z]),
X + Y < Z,
label([X, Y, Z])

See Constraints for the full API.


Why not allow free intermingling of Python and logic namespaces?

Three main reasons:

  1. Ambiguity. It is impossible at compile time to distinguish a Python global from an atom without tracking all imports. Old compiled code could silently become wrong when a new name is imported. With explicit ++ / -- escaping, the boundary is always visible.

  2. Term representation efficiency. A compound term is a plain tuple cell (functor, *args) and an atom is the interned str, so neither needs a wrapper. Allowing arbitrary Python objects as functors would require a boxing wrapper, which is heavier.

  3. Logic variables must be visually distinct. They are declared implicitly, work differently from Python names, and their bindings are reverted on backtracking. A clear syntactic marker — a capital initial (X, FOO, Foo) or a leading underscore (_x) — avoids confusion without requiring explicit declare statements.

The escape mechanisms (--, ++) cover all cases where interop is genuinely needed. Explicit is better than implicit.



Syntax cheat sheet

# Variables (ALL-CAPS preferred; leading-underscore and TitleCase also valid)
X, HEAD, REST          # ALL-CAPS logic variables
_x, _head, _rest       # leading-underscore style (also valid)
Total                  # TitleCase (the ISO spelling) — never in functor position
_                      # anonymous variable (always unifies, stores nothing)

# Atoms (symbols) → the interned Python str
red, café              # bare identifiers: must be declared (-module/-private/-import_from)
'hello world', 'not'   # single quotes: an atom in any spelling, no declaration needed

# Lists
[]                     # empty list
[a, 1, X]              # simple list
[FIRST, *REST]         # head/tail

# Strings ("…" is a string by default: a list of chars)
"hello"                # the carrier ('$chars', 'hello')
b"hi"                  # list of byte codes [104, 105]

# Dicts (DictTerm — keys ground, values may be Vars)
{'x': 1, 'y': 2}              # ground dict (atom keys)
{'x': X, 'y': Y}              # dict with variable values
{'addr': {'city': C}}         # nested dict

# Sets (SetTerm — elements must be ground)
{1, 2, 3}                     # set of integers
{'red', 'green', 'blue'}      # set of atoms

# Unification (bare `is` never evaluates)
X is Y,                # unify
X is 1 + 2,            # X = the term 1 + 2, not 3
'='(X, Y),             # ISO =/2, the same unification
X is not Y,            # dif constraint (must stay different)
not (X is Y),          # immediate check (don't unify right now)

# Arithmetic / CLP(ℤ) constraints
(N == X + 1),          # arithmetic constraint
X == Y,                # arithmetic equality constraint
X != Y,                # arithmetic disequality constraint
X < Y,                 # less-than constraint
X <= Y,                # less-or-equal constraint
eval_(X + 1, N),       # eager arithmetic evaluation
'is'(N, X + 1),        # ISO is/2 (quoted operators follow Scryer: see operators.md)
in_domain(X, 1, 10),    # post integer domain
all_different([X,Y,Z]), # pairwise disequality
label([X, Y, Z]),      # enumerate solutions (first-fail)
'=='(X, Y),            # structural identity (ISO ==/2); also structural_eq(X, Y)

# Rules and facts
head(X) <- goal(X)            # single-call body (no parens needed)
head(X) <- (X > 0, goal(X))   # operator or multi-goal body (parens required)
fact(1),                      # fact (trivially true)
rule >> list_description      # DCG rule

# Goals
goal(A, B),            # compound goal
not goal,              # negation as failure
assertz(fact(1)),      # add a fact at runtime (fact/1 declared -dynamic)
retract(fact(1)),      # retract first matching clause
p(x=1),                # SyntaxError: keyword-argument terms are refused

# Module qualification
utils.double(X, Y),    # qualified call (after -import_module(utils))

# Escaping
++python_expr          # a Python value inside logic code
--logic_term           # a logic term inside Python code (run it in goal position)
~~python_expr          # capture as a Python ast node

# Lambdas (anonymous clauses)
call_goal((X <- (R == X + 1)), 5)                             # R = 6
call_goal(((X, Y) <- (R == X + Y)), A, B)                 # multi-param

# Meta-predicates
findall(X, in_(X, [1,2,3]), BAG),          # BAG = [1,2,3]
bagof(X, in_(X, LIST), BAG),              # fails if LIST empty
setof(X, in_(X, XS), BAG),               # deduplicates
forall(in_(X, NS), X > 0),               # universal quantification
call(GOAL, ARG1),                          # call/2 (alias for call_goal/2)

# F-strings — logic variables auto-deref at search time
writeln_text(f"Hello, {NAME}!"),            # prints bound value of NAME
writeln_text(f"{X:.2f}"),                   # format specs work
write_text_to_string(f"{X} and {Y}", S),    # capture as string

# DCGs — >> defines grammar rules with difference lists
greeting >> (['hello', 'world']),          # terminal sequence
sentence >> (noun_phrase, verb_phrase),    # non-terminal chain
digit(D) >> ([D], {D >= 0}, {D <= 9}),    # args + inline goals
letter >> (['a'] or ['b'] or ['c']),      # disjunction
not_a >> (not ['a'], [X]),                # negation
(peek(T), [T]) >> ([T]),                  # pushback/semicontext
phrase(greeting, ['hello', 'world']),     # phrase/2 — must consume all
phrase(digit(D), [3], REST),              # phrase/3 — partial parse

# EDCGs — Extended DCGs (multiple named accumulators + passed args)
-edcg_acc(counter, X, IN, OUT, {OUT == IN + X})  # declare accumulator
-edcg_pass(config)                                # declare passed arg
-edcg_pred(inc, 0, [counter])                     # declare pred's hidden args
inc >> ([1] // counter)                   # [value] // acc — push to accumulator
get(V) >> (counter / V)                   # acc / Var — read current value
scaled >> (scale / S, [S] // counter)     # pass / Var — read passed arg

# Python interop — ++() evaluates Python at search time
N is ++len(L),                       # call Python builtin
R is ++S.upper(),                    # method call on deref'd var
R is ++(X + 1),                      # Python arithmetic
R is ++L[0],                         # subscript access
R is ++D[K],                         # dict access
++print(X),                          # side-effect goal

# Higher-order list predicates
maplist(GOAL, [1, 2, 3]),                   # check GOAL on each element
maplist(GOAL, XS, YS),                      # map GOAL(X, Y) over list
include(GOAL, LIST, KEPT),                   # keep where GOAL succeeds
exclude(GOAL, LIST, REMOVED),               # keep where GOAL fails
foldl(GOAL, LIST, ACC0, RESULT),         # left fold with GOAL(Elem, Acc, Next)
list_item(INDEX, LIST, ELEM),                # 0-based index access
in_check(ELEM, LIST),                        # deterministic membership check
unpack(TERM, LIST),                         # decompose/construct term

When the syntax is wrong

.seam is Python surface syntax, so a malformed clause is reported by CPython's parser — and a parser reports where it gave up, not where the mistake is. For a rule body that is almost always the closing ), one or more lines below the defect. The loader therefore prints the source itself (clausal/syntax_diagnostics.py, hooked into the parse in clausal/import_hook.py):

invalid syntax (m.seam, line 6)
    3 | f(X) <- (
    4 |     X > 1,
    5 |     Y is
      |         ^ `is` has no right-hand side
    6 | )
      | ^ parse gave up here
  -> complete the expression, or delete the goal — a Clausal goal cannot end
     on an operator.

Read it bottom-up: line 6 is where the parser stopped and is marked as such; line 5 is the line to edit. Up to three preceding lines are shown, plus the enclosing clause head when it falls outside that window (with ...N lines omitted), so a long body cannot bury the diagnosis.

The construct is named where it can be inferred with confidence — a dangling operator (is, +, >, =) with no right-hand side, a <- whose multi-goal body is not parenthesised, a goal not followed by ,, an unclosed (, an unterminated string, and the two Prolog habits head :- body and a trailing ).. When nothing can be inferred the source and caret are still shown and no guess is offered.

Everything CPython set is preserved: the message's first line is its own, and msg, lineno, offset, text and the exception class (SyntaxError, IndentationError, TabError) are unchanged. Genuine .py files are not touched — their errors are Python's to report.


See also: Tutorial — hands-on introduction to Clausal Prolog · Predicates & Rules — clause forms, dispatch, and guards · Builtins — full predicate reference.