Semantic linter (CI)
Your team’s decisions, patterns, and gotchas live in .lore.md, a version-controlled record of how this codebase is supposed to work. The semantic linter reads that record and, on every pull request, flags changes that appear to contradict it.
It is a judge, not a rule engine. Instead of matching regexes, it asks an LLM whether a specific diff hunk conflicts with a specific documented invariant, and surfaces the ones that do as GitHub annotations. It is advisory by default: findings and health failures do not fail the build. A human decides what to do with each finding. A repository can deliberately promote calibrated rules to gate mode, which fails closed when the run is inconclusive.
✓ no suspected invariant violations among selected candidates (45 hunks × 67 invariants → 20 candidates → 20 judge calls)What it is good for
Section titled “What it is good for”- Surfacing the “we decided not to do this” cases that a reviewer would catch only if they happened to remember the original decision.
- Turning tribal knowledge in
.lore.mdinto a check that runs whether or not the person who wrote the rule is reviewing. - Doing this cheaply. Most hunk/invariant pairs are eliminated before any model is called (see How it works).
It is not a replacement for tests, type checking, or a linter. It has no ground truth; promote only narrowly scoped, calibrated rules and keep the rest advisory.
Quick start (GitHub Actions)
Section titled “Quick start (GitHub Actions)”The repository ships a reusable composite action and a reference workflow. It uses the workflow’s GitHub token with GitHub Copilot by default, or you can configure the custom credential and model pair described below.
Add .github/workflows/semantic-linter.yml:
name: Semantic linter
on: pull_request_target: types: [opened, synchronize, reopened]
concurrency: group: semantic-linter-${{ github.event.pull_request.number }} cancel-in-progress: true
permissions: actions: read contents: read pull-requests: read copilot-requests: write
jobs: lint: runs-on: ubuntu-latest timeout-minutes: 25 # Set this protected repository variable only after calibration. continue-on-error: ${{ vars.LORE_SEMANTIC_LINT_GATE != 'true' }} steps: - uses: actions/checkout@v6 with: # Execute only trusted base code. The PR head is fetched as diff data. ref: ${{ github.event.pull_request.base.sha }} fetch-depth: 1 - name: Fetch exact PR head for diffing env: PR_HEAD_SHA: ${{ github.event.pull_request.head.sha }} PR_NUMBER: ${{ github.event.pull_request.number }} run: | set -euo pipefail sha_re='^[0-9a-fA-F]{40}$' [[ "$PR_HEAD_SHA" =~ $sha_re ]] || exit 1 [[ "$PR_NUMBER" =~ ^[1-9][0-9]*$ ]] || exit 1 git fetch --no-tags --depth=1 origin \ "+refs/pull/${PR_NUMBER}/head:refs/remotes/origin/lore-pr-head" test "$(git rev-parse refs/remotes/origin/lore-pr-head)" = "$PR_HEAD_SHA" - uses: pnpm/action-setup@v6 - uses: actions/setup-node@v6 with: node-version: "24" - run: pnpm install --frozen-lockfile - run: pnpm --filter @loreai/gateway run bundle - name: Run Lore semantic linter uses: ./.github/actions/lint with: base: ${{ github.event.pull_request.base.sha }} head: ${{ github.event.pull_request.head.sha }} lore-command: "node packages/gateway/dist/bin.cjs" model: ${{ secrets.LORE_WORKER_API_KEY != '' && vars.LORE_INVARIANT_MODEL != '' && vars.LORE_INVARIANT_MODEL || 'github-copilot/gpt-5.6-luna' }} worker-api-key: ${{ secrets.LORE_WORKER_API_KEY != '' && vars.LORE_INVARIANT_MODEL != '' && secrets.LORE_WORKER_API_KEY || '' }} github-token: ${{ secrets.LORE_WORKER_API_KEY != '' && vars.LORE_INVARIANT_MODEL != '' && '' || github.token }} gate: ${{ vars.LORE_SEMANTIC_LINT_GATE == 'true' }}Open a PR and the check runs, posting any suspected contradictions as annotations plus a job summary. The reference workflow passes a 20-minute overall deadline and a 90-second per-candidate timeout, leaving five minutes for report publication and gateway shutdown.
PR runs restore a derived invariant database but never write it. Copy the repository’s semantic-linter-cache.yml too: it primes that cache on trusted main changes, including commits that change only .lore.md, avoiding forbidden cache-save attempts from pull_request_target runs.
Choosing a judge model and credential
Section titled “Choosing a judge model and credential”The credential and model are selected as a pair to prevent sending one provider’s key to another provider.
Credential (worker-api-key or github-token):
- Default. The reference workflow passes
github.tokento the action’s loopback-only official Copilot SDK bridge. The SDK resolves the Actions installation token and its billing identity;copilot-requests: writegrants inference access. The token is not forwarded directly toapi.githubcopilot.com. - Custom provider. Set
LORE_WORKER_API_KEYandLORE_INVARIANT_MODELtogether. The custom pair takes precedence over the workflow token.
Model (the model input, provider/id):
| Situation | Model used |
|---|---|
LORE_WORKER_API_KEY and LORE_INVARIANT_MODEL are set |
the variable value, authenticated by the dedicated key |
| Only one custom override is set | github-copilot/gpt-5.6-luna, authenticated by the workflow token |
| Neither custom override is set | github-copilot/gpt-5.6-luna, authenticated by the workflow token |
Why use the Copilot SDK bridge?
Section titled “Why use the Copilot SDK bridge?”Lore already knows how to call GitHub Copilot’s Chat Completions and Responses endpoints. The bridge exists for credential handling, not model transport.
A bearer credential already accepted by the Copilot model API, such as the user OAuth credential stored by OpenCode, can use Lore’s direct github-copilot provider path. The workflow’s github.token is different: it is a short-lived GitHub App installation token. copilot-requests: write authorizes Copilot use through supported GitHub tooling, but does not turn that installation token into a supported bearer credential for direct requests to api.githubcopilot.com. Direct forwarding was tried before the bridge was introduced and produced HTTP 400 responses even after the request body was corrected.
The official Copilot runtime handles GitHub’s authentication, entitlement, repository policy, endpoint selection, and billing attribution for the installation token. The SDK gives the action structured control over that runtime: it creates tool-disabled sessions with Lore’s exact judge prompt, propagates cancellation, and performs bounded cleanup. Lore still owns candidate selection, model routing, retries, verdict parsing, and reporting; the action token is never forwarded by Lore to the model endpoint.
The SDK and CLI are installed from the isolated, integrity-locked .github/actions/lint/copilot-sdk package so this optional CI path does not add a platform-specific Copilot binary to every Lore installation. That runtime could be removed if the action required a bearer credential already accepted by the Copilot model API instead, or if GitHub published a supported direct API contract for Actions installation tokens. Reimplementing the CLI’s private authentication and billing behavior in Lore would be brittle and unsupported.
The github-token bridge currently accepts only github-copilot/gpt-5.6-*. This is a Lore bridge implementation constraint, not an official SDK model limitation: the SDK supports every model available in Copilot CLI. Supporting other Copilot models requires updating the bridge and its local gateway route together. Use worker-api-key until then for a different provider or direct Copilot credential.
For a personally owned repository, Copilot usage is billed to the repository owner’s Copilot seat. Organization-owned repositories must enable Allow use of Copilot CLI billed to the organization in their Copilot policy settings. These billing and policy requirements are separate from the workflow permission.
How it works
Section titled “How it works”The check is a three-stage funnel designed so the expensive stage runs as rarely as possible:
- Changed-files gate. Only files touched by the PR are considered.
- Embedding cosine prefilter (free, local ONNX). Every diff hunk is embedded and matched against the invariant embeddings. The vast majority of hunk/invariant pairs are semantically unrelated and dropped here. A large PR can generate thousands of pairs, of which only a handful survive.
- LLM judge. The surviving candidates (capped at 20 per run) are sent to the judge one pair at a time: does this hunk contradict this invariant? Only these calls cost tokens.
The funnel line in the report (N hunks × M invariants → C candidates → J judge calls) shows how aggressively each stage narrowed the work.
On a wide PR, the cap means a complete, clean report says none of the selected candidates looked contradictory; it is not proof that every possible hunk/invariant pair was examined. Split broad changes before relying on an enforced rule.
Where the invariants come from
Section titled “Where the invariants come from”In CI there is no local Lore database, so the action derives one from the committed .lore.md: it imports the plaintext entries and embeds them in-process. That derivation is cached with actions/cache, keyed on the judge model, the onnxruntime-node version, and the .lore.md content hash, so a stale embedding space is never silently reused (embedding drift would quietly rot recall). The cache only rebuilds when the knowledge or the embedding stack changes.
Which invariants are eligible
Section titled “Which invariants are eligible”Not every .lore.md entry is a candidate. The check only considers prescriptive invariants: entries that state a rule (“always…”, “never…”) that a code change could actually contradict. Descriptive facts about workflow, sessions, or personal preferences are skipped, because a spurious flag there is pure noise. Enumeration-style invariants (lists that are expected to grow) are surfaced at most as advisory notes, since reordering or extending a list is legitimate drift, not a violation.
Tuning
Section titled “Tuning”Reasoning effort
Section titled “Reasoning effort”--effort (or the invariantCheck.effort config key) is a cost/depth dial for the judge on reasoning-capable models. It accepts off | low | medium | high | xhigh and defaults to off.
- On a reasoning model, higher effort spends more tokens reasoning about each hunk/invariant pair, which helps when subtle contradictions are being missed.
- On a non-reasoning model, it is ignored.
Set it per-repo in .lore.json:
{ "invariantCheck": { "effort": "medium" }}Or per-run via the action’s effort input, or the --effort CLI flag. The flag overrides the config value.
Running it locally
Section titled “Running it locally”The same check is available from the CLI, which is the fastest way to try it against a real range before wiring up CI:
lore lint --base <sha> --head <sha>With no arguments it auto-detects the range (the current branch against its base). Useful flags:
--model <provider/id>sweeps a specific judge model.--effort <level>sets reasoning effort, as above.--project <path>checks a different working tree.--jsonemits the versioned machine-readable report on stdout for local tooling.--report-file <path>atomically writes a validated, versioned JSON report. CI uses this owned channel instead of redirecting stdout.--deadline-ms <ms>bounds the overall run (default:1200000).--candidate-timeout-ms <ms>bounds each selected judge candidate (default:90000).
The CLI exit contract is:
| Exit | Meaning |
|---|---|
0 |
Complete, non-blocking report |
1 |
Argument/usage failure before report setup |
2 |
Complete gate-mode report with blocking findings |
3 |
Partial or failed runtime report |
The action always consumes and validates --report-file, regardless of the CLI exit. Missing, malformed, wrong-version, or internally inconsistent reports are health failures. Advisory actions show those failures but remain non-blocking; gate actions fail closed.
Report health
Section titled “Report health”The report distinguishes complete, partial, and failed runs. It records health for range resolution, diff parsing, invariant loading, invariant vectors, hunk vectors, and the judge. Every selected candidate is recorded as resolved, unresolved, or not attempted, and the report validator checks that candidate states and semantic/transport attempt totals match the funnel counters.
“No suspected invariant violations” is shown only for a complete report with zero findings. Partial and failed reports remain visibly inconclusive even when they contain no findings.
Enforcement tiers
Section titled “Enforcement tiers”The linter ships advisory-only: findings are surfaced, nothing blocks. This is intentional. A probabilistic judge is right to inform a human but wrong to gate on until a team has watched its false-positive rate on their own repo.
A graduated ladder is designed above advisory:
- advisory: a note; never fails a build. (Shipped, the default.)
- soft: an overridable gate. A finding blocks unless the PR author adds a
lore-override: <invariant> — <reason>trailer to a commit in the range. - strict: a hard gate that cannot be overridden.
An invariant only escalates past advisory when its author explicitly opts it in on its marker. Enumeration invariants are always capped at advisory regardless. The --gate flag (and the action’s gate input) is the switch that makes soft/strict findings blocking.
<!-- lore:019e18ec-e328-76c4-9c3c-09dbe8d51c6c enforce:soft -->* **SQLite driver boundary**: `node:sqlite` must only be imported by packages/core/src/db.ts.Use enforce:soft, enforce:strict, or enforce:off; omit it for advisory. Lore preserves the marker on export and imports it into the entry metadata. soft can be overridden with a lore-override: <invariant> — <reason> commit trailer; strict cannot.
Promoting a rule to a gate
Section titled “Promoting a rule to a gate”- Leave a new rule advisory and review its findings across several representative PRs, including refactors and test-only changes.
- Make the rule narrow, prescriptive, and anchored to a file path or symbol; never promote a broad architectural summary or an enumeration.
- Add
enforce:softfirst. After confirming that its findings are actionable, set the protected repository variableLORE_SEMANTIC_LINT_GATEtotrue. - Use
strictonly for a small rule that has already demonstrated a negligible false-positive rate. A partial or failed run blocks in gate mode by design.
The supplied workflow reads that variable from the trusted base repository. With the variable absent or anything other than true, it remains advisory; this makes rollout reversible without accepting a PR-controlled gate switch.