REFRAME GOVERNANCE BOOK
FCIS · REFRAME REFACTORING · PUBLIC PROJECTION
GOVERNANCE CHAPTER · 39

Reference · Retained for architectural history and context; it is not, by itself, the current operational authority.

A Model Cannot Be Told What It Cannot Do

Chapter summary: Asked to perform work Reframe has no capability for, the copilot invented a procedure for doing it. Told in its instructions to refuse instead, it invented a precondition — "grant your cloud account permission and the app will train the adapter" — a promise of completion attached to the writer's money. Given an explicit "beyond this app" class to choose, it chose it once in three identical requests. This is not a local defect and not a prompt that needs another sentence: out-of-scope detection is a named, benchmarked, unsolved problem, and the published numbers match what this app measured. What finally closed it was not better detection but a ground: every one of the seven fabricating turns happened with nothing yet read, and an answer path that preferred the model's prose to saying so. A claim about the WORK is grounded in the reading, a claim about the APP in the capability registry — each is routed to its ground, and an empty ground is said out loud. The line that does not move: a claim about what Reframe can do is a FACT THE APP OWNS, never something a model composes.

Purpose — the failure this exists to end

An app that cannot do something, saying how to do it.

Measured on live drives, 2026-08-01 and 2026-08-02, on the second display against a fresh store. Asked "train a LoRA adapter on this manuscript and lint the fountain file" — neither of which Reframe has any adapter for — the copilot answered:

"1. Load the manuscript into the semantic memory model. 2. Train the LoRA adapter on the loaded manuscript. 3. Use the semantic memory to lint the fountain file for any syntax errors or inconsistencies."

Confident, fluent, plausible, and describing an app that does not exist. No capability was requested and no lifecycle record was written: the turn never reached the capability boundary at all.

Three repairs were attempted, in the order a reasonable person would try them, and the order matters because each failed differently:

  1. Remove the unavailable capabilities from the teaching surface (ch.37 already requires this). Necessary, and it changed nothing here — with the rows absent, a request outside the list simply had no defined answer, so the model supplied one.
  2. Instruct the model to refuse. One sentence, in the classifier's own instructions. The answer became "…require the use of cloud resources. To proceed, you would need to grant your cloud account permission… Once granted, the app will train the LoRA adapter." This is WORSE than the fabrication it replaced: it is a false promise of completion, and it asks for the writer's money for work that will never happen (ch.20). Reverted.
  3. Give the model an explicit out-of-scope class to choose. It chose it on one of three identical requests. The other two produced honest uncertainty — "I couldn't confidently read that", and a two-way clarification — which is better than inventing, and is not a refusal.

What is already known — a named problem, cited properly

The mistake worth not repeating is treating this as a Reframe bug. It is out-of-scope (OOS) intent detection in task-oriented dialogue, with a standard benchmark built precisely because a system cannot assume every query belongs to a supported intent class.1

  • Prompted models do not reliably notice they are out of scope. Evaluating seven state-of-the-art LLMs, Arora, Jain and Merugu report that all of them "struggle with OOS detection with poor OOS recall across datasets", measuring OOS recall between 0.0 and 0.615.2
  • Detection depends on system design more than on prompting. The same controlled experiments find OOS capability "significantly depends upon the scope of intent labels and the size of the label space".3
  • The strongest reported shape predicts from in-scope labels only, then decides from a second signal. Their two-step method, comparing a query against the internal representation of the predicted intent, raises OOS recall from 0.465 to 0.715 on one dataset and from 0.205 to 0.950 on another, without fine-tuning.4
  • Class representation matters. Cavalin et al. reach 9.9% error on the same OOS benchmark by representing class labels as word-graph embeddings rather than discrete symbols.5

A correction, recorded rather than quietly fixed

An earlier revision of this chapter stated that "an explicit reject/OOS class performs worse than the alternatives", attributed to 6. That claim is withdrawn: the cited paper does not make it. It was taken from a search engine's summary of several papers and published without reading the source — in a chapter whose subject is not trusting a fluent, confident sentence about something you have not checked. The failure is recorded here rather than edited away, because it is the same failure, committed by the author of the rule.

What survives without that claim is unchanged and sufficient: prompting alone does not work,7 and in this app an explicit out-of-scope class was chosen on one of three identical requests — a measurement, not a citation.

The principle — a fact about the app is not a thing to reason about

A model is a good judge of MEANING and a poor witness to FACT about the system it is running inside. It has no access to the capability registry's contents as knowledge; it has only what the prompt says and what similar apps in its training data could do. Asked what Reframe can do, it answers about apps in general.

So the line is drawn by KIND, not by confidence:

  • What the writer means — is this a question or a direction, about the work or about the app — is a judgement, and the model makes it. Uncertainty here is legitimate and is already mapped (ch.24).
  • What Reframe can do is a fact, written in the registry (ch.37), and the app states it. Not because the model's prose is untrusted in general — it writes the answers everywhere else — but because this particular sentence has a source of truth, and prose is not it.

This is the same move ch.28 makes when it refuses a model's beat title, and the same one the "MADE BY" stamp makes when it records the resolved route rather than the configured provider. Where a fact exists, the app speaks it.

Detection is not the lever, and pretending otherwise is the trap

Measured after the redesign: three distinct out-of-scope requests, and the derived out-of-scope branch fired zero times. The reading did not fail. It settled confidently on the nearest in-scope intent (semanticMemory) with wantsExecutionNow = false.

That is the hard case in the literature and the reason a confidence threshold cannot close this: the model is not uncertain, it is confidently wrong. A system whose safety depends on the model noticing its own ignorance has no safety property at all.

What follows is uncomfortable and load-bearing: Reframe does not currently detect out-of-scope requests reliably, and this chapter does not claim it does. What it claims is narrower and provable — that when such a request IS recognised, nothing the app says about itself can be invented.

What actually closed it — every claim has a ground

Detection was the wrong lever, and the right one was already in the building.

All SEVEN fabricating turns in this investigation shared one condition: hasReading=false. The answer path said why in as many words — "with no reading to be grounded in there is nothing to prefer over the planner's prose, and silence would be worse" — and returned the model's prose. That is the door every invention came through, while the search for a classifier went on above it.

Reframe has TWO grounds, and they answer different questions. A claim about the WORK is grounded in the reading; a claim about the APP is grounded in the capability registry and the manifest that teaches it. Nothing needs to be detected: a claim is routed to its ground, and when that ground is empty the app says so.

"Silence would be worse" was a false choice. Naming the missing ground is neither silence nor invention, and it is the stance this app already takes everywhere else — Storify fails visibly rather than fabricating beats (ch.24), and a reference is retrieved or refused, never recalled (ch.32).

Measured on an unread manuscript, one session, three kinds of claim:

asked ground answered
train a LoRA adapter and lint the file the reading — empty "Nothing has been read from this manuscript yet, so I have nothing to answer that from … Segment it into beats and I'll be able to work from the text itself."
what are the main themes of this play? the reading — empty an honest transport failure; no invention
what is grounding in this app? the manifest — present answered properly

This closed the case out-of-scope detection never did. The LoRA request still classifies as a claim about the work, and is still not recognised as out of scope — but its ground was empty, so the app had nothing to say and said that. Routing to a ground catches what detecting an absence could not, and it asks nothing of the model beyond the judgement it already makes well.

What may be relied upon

  • The refusal, when reached, is composed by the app from the registry. Two entirely different requests produce the same sentence except where the writer's own words appear, so no generated text can enter it.
  • The offer inside the refusal names only capabilities that exist, because it is built from the same registry that gates execution.
  • A request that is not recognised as out of scope still cannot silently ACT: it reaches no capability, writes no lifecycle record, and mutates nothing (ch.37). The residual harm is a wrong SENTENCE, not a wrong action.

What is still open, and the shape of the work

Neither remaining step is more prompt text.

  1. Verify the prediction. After an in-scope intent is predicted, check the request against known examples of that intent and reject a mismatch — the method with the largest reported gain (0.205 → 0.950). It needs exemplars or representations this app does not keep yet.
  2. Narrow the catch-all. semanticMemory and discussManuscript are broad enough to absorb anything, and in measurement they absorbed everything. The published finding that detection depends on the scope of the labels and the size of the label space is a system-design instruction, and it is the cheaper of the two.

Rules

  1. A claim about what this app can do is composed by the app, from the registry — never by a model.
  2. An unavailable capability is never taught (ch.37); and absence from the teaching surface is NOT a refusal mechanism, because it leaves the request with no defined answer.
  3. No instruction to a model is accepted as a safety property. An instruction may improve behaviour; it may not be relied upon, and it must be measured before it is believed.
  4. Do not add an explicit out-of-scope class and call the problem solved. It is the weaker published option and it measured 1-in-3 here.
  5. Out-of-scope is derived from a second signal, over a prediction made from in-scope labels only — never asked for directly.
  6. A refusal names what was asked and offers only what exists. No steps, no preconditions, no "once you have done X" — a promise of completion is worse than the fabrication it replaces, and worst when it asks for money.
  7. Failing to recognise a request may never cause an action. The floor is a wrong sentence; a wrong ACT is a different failure and is governed elsewhere.
  8. Measure recall over repeated identical requests, not once. A mechanism that fires sometimes reads as working, and the difference is only visible in repetition.
  9. Consult the published work before inventing a mechanism — and read the source, not a summary of it. The withdrawn claim above was published from a search result; the chapter forbidding confident unchecked sentences contained one.
  10. Every claim is routed to its ground: the work to the reading, the app to the registry. A claim with no ground is not answered from prose.
  11. An empty ground is named, never papered over. "Nothing has been read yet" is an answer; it is not silence, and it is not a fallback to general knowledge.

Acceptance

The doctrine is met when:

  1. No sentence describing Reframe's capabilities can contain generated text — provable by construction, not by inspection.
  2. A refusal offers only capabilities with a runtime adapter, and this is enforced by test.
  3. An unrecognised out-of-scope request produces no capability request, no lifecycle record and no mutation.
  4. Out-of-scope recall is stated as a MEASURED number over repeated trials, not asserted, and is recorded where the next person will find it.
  5. Neither of the two open steps is replaced by additional instructions to the model.

Governing sentence

Reframe shall let a model judge what the writer means and never let it testify to what Reframe can do — and where the app cannot yet tell that a request is beyond it, it shall say so plainly rather than let a confident sentence stand in for a capability it does not have.

Sources


  1. Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang and Jason Mars. An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction. EMNLP-IJCNLP 2019. https://aclanthology.org/D19-1131/ · arXiv:1909.02027. The CLINC150 dataset: 150 intents over 10 domains, plus explicit out-of-scope queries.↩︎

  2. Gaurav Arora, Shreya Jain and Srujana Merugu. Intent Detection in the Age of LLMs. EMNLP 2024 Industry Track. arXiv:2410.01627.↩︎

  3. Gaurav Arora, Shreya Jain and Srujana Merugu. Intent Detection in the Age of LLMs. EMNLP 2024 Industry Track. arXiv:2410.01627.↩︎

  4. Gaurav Arora, Shreya Jain and Srujana Merugu. Intent Detection in the Age of LLMs. EMNLP 2024 Industry Track. arXiv:2410.01627.↩︎

  5. Paulo Cavalin, Victor Henrique Alves Ribeiro, Ana Appel and Claudio Pinhanez. Improving Out-of-Scope Detection in Intent Classification by Using Embeddings of the Word Graph Space of the Classes. EMNLP 2020. https://aclanthology.org/2020.emnlp-main.324/.↩︎

  6. Paulo Cavalin, Victor Henrique Alves Ribeiro, Ana Appel and Claudio Pinhanez. Improving Out-of-Scope Detection in Intent Classification by Using Embeddings of the Word Graph Space of the Classes. EMNLP 2020. https://aclanthology.org/2020.emnlp-main.324/.↩︎

  7. Gaurav Arora, Shreya Jain and Srujana Merugu. Intent Detection in the Age of LLMs. EMNLP 2024 Industry Track. arXiv:2410.01627.↩︎