In a structural analysis we reviewed in September, the engine reproduced its results byte for byte and passed the global force and moment balance checks. The calculation was ready to support further analysis. Structural design still required a separate decision about the assumptions the model had used.
That was a useful stopping point. We had evidence that the computation could be reproduced, with the expected balance of loads and reactions. We also had assumptions whose status was still provisional. Keeping those facts separate let us close the computational work within its stated scope and leave the design decision open.
It also made me look more carefully at the setup around the model. We use knowledge graphs, skills, and engines together in our engineering work. Each contributes something that becomes visible at a different point in the result.
What each component contributes
By an engine, I mean maintained software that performs a defined operation on explicit inputs. It might calculate a structural response, read geometry from a design file, or produce an artifact the next tool can consume. It has a scope, and its output can be inspected separately from the conversation that requested it.
A knowledge graph keeps the conditions around a technical criterion connected. A value carries the assumptions that make it applicable, the other criteria that constrain it, and a path back to its source. That is the kind of graph I described in A knowledge base is a graph, not a folder. Its value comes from the relationships it preserves. Those relationships still need to be checked.
A skill is a reusable procedure the agent can consult when performing a task. It carries the practical knowledge needed to use the tools and recognize a usable result. A skill can include scripts or point to an engine. Reading the procedure and running the software are two events we need evidence for.
These are responsibilities with some overlap. An engine can implement a rule from the graph. A skill may include a calculation. The distinction helps us understand which part of a result we are relying on.
| Component | Main contribution | Evidence we can inspect |
|---|---|---|
| Knowledge graph | Conditions and relationships behind a criterion | Applicable sources, assumptions, and the reasoning supported by them |
| Skill | A reusable way to carry out the task | The procedure used and the artifacts it calls for |
| Engine | Computation or transformation over defined inputs | Actual inputs, outputs, and the checks relevant to that operation |
| Model | Interpreting the request and coordinating the work | The actions taken and the explanation tied to the artifacts |
| Responsible professional | Accepting the engineering decision within its scope | The decision and the evidence on which it rests |
The model has a substantial job here. It has to recognize what the task requires, consult the relevant material, use the procedure, and interpret what comes back. The surrounding software gives us places to inspect that work.
The connection shows up in the output
The road-design case I wrote about in Give the agent the real source of truth ended with 26 relief culverts placed in the design software. We checked the calculation against a structure already placed by hand. The output also had to survive being loaded into the target software.
That case connected several different claims. The geometry had to come from the computed road model. The calculation had to place the ends correctly. The generated artifact had to be readable by the application. Once the solution worked, we made the procedure reusable through a skill.
An engine gives that computational part a stable place to live. We can inspect what went in, compare what came out, and maintain the routine independently of the wording of the next request. Reproducibility has a scope too: inputs, versions, and relevant numerical settings have to be controlled. Some operations involve numerical approximations, and some tools include probabilistic components.
The graph contributes at the point where a technically valid operation needs a reason to apply to this case. Our earlier geometry post describes a rule we had encoded too conservatively and then corrected. The structure helped expose the problem; the stored criterion itself needed repair.
That is why the output matters so much. A persuasive explanation, a procedure available in context, and a successful software run each support a particular claim. The resulting artifact lets us check how those claims connect.
A successful run leaves a specific decision
In the structural analysis at the start of this post, the reproduced results and balance checks supported the elastic-analysis scope we had reviewed. They gave us a basis for continuing the work. The assumptions used for that analysis still needed separate consideration before they could govern a design decision.
We kept the numerical result as evidence for that limited scope. The decision about its engineering use remained explicit. It would have been easy to compress the whole thing into “the engine passed.” That would have hidden the question the checks had actually answered.
This separation also makes a correction more useful. If the source criterion is wrong, that criterion needs attention. If a routine transforms the inputs incorrectly, the software needs repair. If the agent skips a required procedure, the gap is in how the procedure was used. The evidence gives us a place to investigate each case.
Some of those corrections become Functional Scars: controls attached to observable situations where an error might recur. Their activations provide evidence about the control’s operation. Whether the task ended correctly remains a separate question about the outcome.
What we can claim from this
I don’t have a controlled benchmark showing how much this combination improves engineering decisions across projects. The examples here support narrower observations: a repeatable calculation with explicit limits, a generated artifact checked in its target software, and a criterion corrected after review.
Building and maintaining the surrounding material takes work. For a simple lookup, ordinary retrieval may be sufficient. We get more value from this arrangement when a task combines conditional technical judgment with computation and a deliverable that can be inspected.
In the reviewed structural case, the computational evidence stayed usable without becoming a design approval. In the road-design case, the inspected output and the reusable procedure both survived the original conversation. Those are the results I want to preserve when the next task arrives.