Laboratory development programme · Stage 1

LydStone Lab: prove it before you deploy it

Independent AI solution audit and qualification laboratory

The laboratory tests which tasks a specific AI configuration can handle, where its sufficiency ends and what resources the result requires.

The first stage of the laboratory development programme is complete: a study using synthetic tasks in a local environment.

Where the laboratory could help

№1

Approved yesterday. Changed today.

An engineer approves a solution. Later, one parameter in a source document changes.

What we test →

№2

The cheapest proposal was incomplete.

One supplier includes installation and testing. Another quotes equipment only. A third substitutes alternatives for part of the specification.

What we test →

The tasks selected for Stage 1 — and why

Stage 1 of the programme covered seven task groups found in work with requests, documents and operational signals.

These tasks can hide an error behind a plausible answer, while correctness can be checked against predefined facts, rules and constraints.

  • W1

    Request clarification and dialogue

    If a key condition is missing, AI should ask for it rather than silently invent it.

    A missing-information clarification scenario was studied.

  • W3

    Extracting data from documents

    The system must find the right fact and distinguish missing information from an invitation to invent it.

    The selected corpus was too easy to distinguish configurations.

  • W4

    Reconciling conflicting sources

    Two convincing sources can disagree. AI must notice the conflict and avoid presenting a guess as a supported conclusion.

    Recognition of irreducible ambiguity was studied.

  • W5

    Inference through a chain of rules

    Several correct steps do not guarantee a correct conclusion: it must follow from the given facts and rules.

    Rule chains with increasing difficulty were studied.

  • W6

    Cause diagnosis

    A plausible explanation of a fault does not establish its cause or justify the next action.

    Measurements were collected, but the difficulty ladder did not reliably distinguish configurations.

  • W7

    Monitoring and signal triage

    Transient noise can be mistaken for a persistent problem and direct attention to the wrong place.

    Signal discrimination under changing transient noise was studied.

  • W9

    Writing under constraints

    Fluent text can omit a required condition or include something explicitly forbidden.

    Compliance with simultaneous required and forbidden conditions was studied.

One synthetic scenario with difficulty levels was tested in each group. Conclusions apply to those scenarios and configurations, not to every task in the group. W3 and W6 do not serve as routing evidence.

Tool use (W2) and retrieval with knowledge-base answers (W8) were separated and deferred: each needs its own evaluation tools and methodology.

The LydStone Lab development programme

The programme aims to make AI selection for a specific process a testable engineering decision, connecting task requirements, evidence of fitness and resource needs.

Stage 1 established the foundation: the CogniScope measurement system, a set of task groups, reproducible experiments and explicit evidence boundaries. The methodology underwent independent audit and re-verification, with reservations retained.

Further development concerns customer-specific examples, broader scenarios and testing under different conditions. Hardware transfer, load and operational reliability each require their own validation; Stage 1 does not establish them.

CogniScope — the programme’s technical core

CogniScope is the AI qualification research system used to build the LydStone Lab evidence base. It connects scenarios, model configurations, reproducible runs, response evaluation, resource measurements and reports traceable to experimental records.

In Stage 1 of the programme, the system was used for local synthetic testing. Further work develops its capabilities; readiness must be established by separate results.

The unit of qualification is configuration × workload × execution profile. Quality is considered alongside latency, VRAM and output tokens.

In the tested scenarios, the same configuration has different sufficiency boundaries. For example, 27B Q4 without reasoning reaches L0 for irreducible ambiguity, L1 for rule-chain inference and L2 for constrained text.

L0–L3 describe difficulty within each scenario. L3+ means the entire tested range passed and no boundary was found within it. It makes no claim about levels outside that range.

Maximum unconditionally qualified levels for four configurations on W4, W5 and W9. 27B Q4 without reasoning reaches L0, L1 and L2. With stronger reasoning it reaches L3+, L2 and L3+. 9B Q4 without reasoning qualifies on none; with reasoning it reaches L1, L2 and L3+ for the 7168 setting on W9. L3+ means no boundary was found within the tested range.

Sufficiency depends on the task. L3+ means no boundary was found within the tested range.

Original Stage 1 figure in Russian. The English explanation preserves its results and limitations.

Open full-size figure

Three findings from Stage 1 of the programme

W4

97.4% of cases passed. Why did the configuration fail?

For irreducible ambiguity, W4, 27B Q4 without reasoning passed L0: 30 of 30 cases, with no fatal errors. Median latency was 884 ms and output length was 51 tokens.

At L1, 38 of 39 cases passed. However, one error violated the predefined zero-fatal-error allowance. Verdict: FAIL (RED).

The case pass rate does not replace acceptance criteria. A good average cannot compensate for an error the process cannot tolerate.

With stronger reasoning, the same 27B Q4 passed the tested range through L3: 15 of 15 cases at L3, median latency of 12,719 ms and 732 output tokens.

27B Q4 without reasoning on W4: L0 passes 30/30 cases with no fatal errors. At L1, 38/39 cases pass but one fatal error causes FAIL under zero allowance.

27B Q4 without reasoning: one fatal error at L1 causes a fail under zero allowance.

Original Stage 1 figure in Russian. The English explanation preserves its results and limitations.

Open full-size figure

W9

The next capability level has a resource cost

For text with required and forbidden conditions, W9, 27B Q4 without reasoning passed L2: 14/14 cases, 1,482 ms, 58 output tokens and 20,809 MB of VRAM.

27B Q4 with stronger reasoning passed L3: 16/16 cases, 15,540 ms, 839 output tokens and 20,807 MB of VRAM.

Moving between these confirmed levels is associated with approximately 10.5× the latency and 14.5× the output tokens. VRAM remained around 20.8 GB.

Difficulty and reasoning mode both change. This is the cost of moving between confirmed levels, not the isolated effect of reasoning.

W9 confirmed L2 without reasoning: 1,482 ms and 58 tokens. Confirmed L3 with stronger reasoning: 15,540 ms and 839 tokens. Approximately 10.5 times the latency and 14.5 times the tokens, with VRAM around 20.8 GB. Difficulty and reasoning mode both change.

Confirmed L2 versus L3. Difficulty and reasoning mode both change; the effect of reasoning is not isolated.

Original Stage 1 figure in Russian. The English explanation preserves its results and limitations.

Open full-size figure

W5

An evidence-based refusal to recommend is a result

For rule-chain inference, W5, no tested configuration achieved unconditional PASS (GREEN) at L3.

Two 27B configurations with stronger reasoning were CONDITIONAL (AMBER): Q4 at 21/22 and Q6 at 21/23. FAIL (RED): 9B Q4 with reasoning at 3/4; 27B Q4 without reasoning at 3/6; 9B Q4 without reasoning at 2/6.

This outcome stays in the report. The next step may be a different decomposition, refined criteria or a new test. Insufficient evidence does not justify declaring a winner.

W5 at L3: 27B Q4 with stronger reasoning is conditional at 21/22; 27B Q6 is conditional at 21/23. 9B Q4 with reasoning fails at 3/4, 27B Q4 without reasoning at 3/6 and 9B Q4 without reasoning at 2/6. No tested configuration is an unconditional pass.

No tested configuration qualifies unconditionally at L3. Conditional and negative outcomes are retained.

Original Stage 1 figure in Russian. The English explanation preserves its results and limitations.

Open full-size figure

What the evidence establishes and where it stops

Stage 1 of the laboratory development programme demonstrated the method’s basic ability to measure configuration sufficiency for a specific AI task in a synthetic local environment.

The methodology underwent independent audit, remediation and independent re-verification. The audit verdicts include reservations; completing the stage does not remove them.

The study is limited to one host, Stage 1 scenarios and pilot statistics. The boundary-search procedure requires further validation. Comparing different model generations and reasoning modes does not isolate the effect of model size.

Stage 1 does not establish production safety, load or concurrent-request behaviour, cloud-versus-local economics, hybrid routing or a full review of a customer architecture.

The insufficiently discriminating W3 and W6 results are retained and do not serve as routing evidence. W2 and W8 are deferred.

Frozen evidence from Stage 1 of the programme: cogniscope-stage1b-final
0dac1339fbaa2cb51741c102bdfa0ac148936d9a

Proposed service

Decision Sprint

The laboratory offers to qualify 1–2 critical AI tasks or nodes before a major pilot, infrastructure purchase or architecture commitment.

We start with the decision that depends on the test. We then record the hypothesis, review how work is divided between AI, deterministic code and people, and agree on configurations, examples and acceptance criteria.

What you provide

  • An architecture hypothesis or workflow graph
  • Anonymised or synthetic examples
  • Data and execution-environment constraints

What you receive

  • A qualification report and a map of critical nodes and criteria
  • A matrix of tested configurations and confirmed or rejected assumptions
  • A resource profile, applicability boundaries and the next validation step

A sprint typically takes 1–2 weeks after the inputs are ready. Scope, schedule and price are agreed after a 30–45-minute scoping conversation.

Architecture ownership and the final decision remain with the client. The sprint is not a safety certification or a substitute for load testing. Cloud-versus-local economics and hybrid routing require separate validation. Testing on real data inside the client perimeter and subsequent operational stages are scoped separately.

Describe the task, the decision you need to make and the error you cannot tolerate. That is enough for an initial message.

Discuss a Decision Sprint

Dmitry Vaneev · LydStone Lab