LydStone Lab · Where the laboratory could help
30 example tasks for LydStone Lab
30 situations where a plausible AI answer is not enough. Choose a familiar example to help frame your task and identify what needs testing.
These are potential directions for testing, not completed LydStone Lab studies or client cases. Results from the first stage of the programme concern individual synthetic tasks and do not establish that all the solutions described here are ready.
Stage 1 results →Start with these eight
Documents, supplier proposals, email, resources, permissions and precise writing.
№1Approved yesterday. Changed today.
Situation
An engineer approves a solution. Later, one parameter in a source document changes.
What we test
Can AI identify which conclusions and approvals now need review? The harder cases hide the change in an appendix and affect several linked documents.
№2The cheapest proposal was incomplete.
Situation
One supplier includes installation and testing. Another quotes equipment only. A third substitutes alternatives for part of the specification.
What we test
We test whether AI can compare proposals on an equivalent scope and flag unknowns before making a recommendation.
№5The most important email never says “urgent”.
Situation
Among dozens of messages, a calm clarification arrives: the client has changed a condition on which the proposal depended.
What we test
We test whether the assistant notices the change in obligations and raises its priority, even when the email is brief and undramatic.
№6An expensive server for one difficult operation.
Situation
A company plans to buy hardware for a local assistant. Which operations actually need a larger model or stronger reasoning?
What we test
Testing individual tasks can inform resource requirements before procurement. Load, concurrency and requirements for the complete system need separate assessment.
№13“Yes, go ahead” — but what was authorized?
Situation
Several options were discussed in a long conversation. A brief approval may refer to drafting a document rather than sending it to the client.
What we test
We test whether the agent can identify what was approved and stop when its authority remains ambiguous.
№18One clause was fixed. Another changed meaning.
Situation
AI is asked to update a deadline while preserving the other terms. In rewriting, a qualification disappears or a related phrase changes.
What we test
We test the precision of a local edit and preservation of everything the user did not authorize it to change.
№23There is a citation. It does not support the claim.
Situation
An assistant answers from a knowledge base and cites a relevant passage. The passage does not establish the specific statement it makes.
What we test
We test the link between each material conclusion and its source, including abstention when the evidence is insufficient.
№26Marketing copy turned an experiment into a proven outcome.
Situation
AI turns a research report into a service page. “Tested on synthetic examples” becomes “proven in business operations”.
What we test
We test whether it preserves evidence status, limitations and the distinction between a result and a proposed capability.
22 more situations
Other tasks in engineering, development, operations and client work.
№3One assistant. Three different limits of trust.
Situation
An assistant is asked to extract parameters, detect contradictions and draft a conclusion. One configuration may handle the first operation but struggle with the others.
What we test
The laboratory tests each node separately before choosing a model for the whole assistant.
№4“No issues found” in an incomplete document set.
Situation
An appendix required for the review is missing. AI may produce a polished conclusion after reading everything available.
What we test
We test whether it distinguishes a lack of detected issues from a lack of evidence needed to reach a conclusion.
№7The sales proposal promised too much.
Situation
AI turns meeting notes into a proposal. “Timing will be agreed after a site survey” quietly becomes “completed in two weeks”.
What we test
We test whether the configuration preserves conditions, qualifications and limits of commitment while writing persuasive copy.
№8The substitute matches every specification but one.
Situation
A supplier offers an alternative component. The main parameters match, but compatibility depends on a condition in another document.
What we test
We test whether AI finds that condition and separates confirmed compatibility from an assumption that a specialist must check.
№9A model update improved answers and broke a working rule.
Situation
The new version writes faster and more clearly, but guesses more often when information is missing. Can it replace the current configuration?
What we test
The laboratory checks whether acceptance criteria still hold on cases that matter to the workflow.
№10A contradiction that is not a contradiction.
Situation
Two documents give different values for one parameter: one describes normal operation, the other an extreme condition.
What we test
We test whether AI accounts for context and avoids false issues that engineers would then have to investigate.
№11Three agents agreed on the same mistake.
Situation
One agent drafts a conclusion, another reviews it, and a third checks quality. All rely on the same misunderstood premise.
What we test
We test whether this design adds reliability on the specific task and which errors pass through every layer.
№12Twenty alerts for one incident.
Situation
A fault generates messages from several systems. The assistant may present them as twenty independent problems or miss a genuinely separate failure.
What we test
We test signal grouping, preservation of meaningful differences and the basis for priority.
№14No response from the system. Retrying may be unsafe.
Situation
An agent sends a request and receives a timeout. The operation may already have succeeded.
What we test
We test whether it distinguishes a confirmed failure from an unknown outcome before creating another request, sending a message or changing a record.
№15Discussed in the meeting. Decided in the minutes.
Situation
AI writes minutes and action items. A suggestion becomes a decision; a tentative date becomes a commitment.
What we test
We test whether it preserves the distinction between a proposal, question, decision, assignment and confirmed owner.
№16One exception changes the whole conclusion.
Situation
A working rule looks simple, but includes an exception that applies only when several conditions coincide.
What we test
We test how far into a chain of rules the configuration can reason reliably before it starts producing plausible answers by analogy.
№17A promising contract the company cannot deliver.
Situation
An assistant screens business opportunities. An attractive budget conceals an incompatible deadline, a mandatory competence or an unsupported way of working.
What we test
We test whether it identifies reasons to decline and asks the necessary questions before recommending the opportunity.
№19A field note became a more certain fact.
Situation
A specialist writes: “Working for now after the restart; we are still investigating the cause.” The summary says: “Fault resolved.”
What we test
We test whether AI preserves the temporary observation and uncertainty when converting field notes into a structured report.
№20The symptom disappeared. The cause remains unknown.
Situation
Readings return to normal after an operator takes action. The assistant declares the repair successful, although the same improvement might have happened without intervention.
What we test
We test whether it distinguishes observed improvement from confirmed removal of the cause and proposes the necessary verification.
№21Nine tasks accepted. The tenth blocks completion.
Situation
AI prepares a project completion report. Most items are closed, but one unfinished item is a condition of overall acceptance.
What we test
We test whether the assistant substitutes a good completion percentage for a mandatory requirement.
№22Correct information attached to the wrong object.
Situation
A system contains similar names for companies, rooms or equipment. AI extracts the data correctly but links it to the wrong record.
What we test
We test entity identification and behavior when there are too few distinguishing details.
№24An external document tries to control the assistant.
Situation
A supplier proposal or incoming email contains instructions to ignore other options, change evaluation criteria or share additional information.
What we test
We test whether the system preserves the boundary between data under review and authorized instructions.
№25A code fix removed a necessary restriction.
Situation
AI fixes a defect by simplifying a check that also limited user permissions. The main scenario works again, but the system has changed beyond the task.
What we test
We test adherence to the scope of the change and preservation of important invariants.
№27Ten packs became ten parts.
Situation
AI converts messages and specifications into a structured order. Every field is filled and the format is valid, but the unit of account has been lost.
What we test
We test semantic accuracy, including quantities, included components and units of measurement.
№28The translation added a promise.
Situation
A Russian proposal cautiously describes a possible result. The English version reads as a guarantee.
What we test
We test preservation of commitments, limitations and degree of certainty when localizing technical and commercial materials.
№29The calendar is free. The event cannot happen.
Situation
An assistant books a venue using only the event start and finish. Setup, dismantling, crew access and equipment preparation are left out.
What we test
We test whether it extracts the right constraints and passes them to a conventional scheduling system.
№30An audio system fits the budget, but not the room.
Situation
An AI adviser confidently recommends a catalogue bundle without asking how and where it will be used.
What we test
We test which clarifications it treats as essential, how it handles unknowns and what supports its compatibility claims.
Does one of these sound familiar?
For the first conversation, describe the workflow, share an example input and identify an unacceptable error. Together we can determine whether 1–2 critical nodes are suitable for a Decision Sprint.