← Back to LydStone Lab research

LydStone Lab · Where the laboratory could help

30 example tasks for LydStone Lab

30 situations where a plausible AI answer is not enough. Choose a familiar example to help frame your task and identify what needs testing.

These are potential directions for testing, not completed LydStone Lab studies or client cases. Results from the first stage of the programme concern individual synthetic tasks and do not establish that all the solutions described here are ready.

Stage 1 results →

Start with these eight

Documents, supplier proposals, email, resources, permissions and precise writing.

№1

Approved yesterday. Changed today.

Situation

An engineer approves a solution. Later, one parameter in a source document changes.

What we test

Can AI identify which conclusions and approvals now need review? The harder cases hide the change in an appendix and affect several linked documents.

№2

The cheapest proposal was incomplete.

Situation

One supplier includes installation and testing. Another quotes equipment only. A third substitutes alternatives for part of the specification.

What we test

We test whether AI can compare proposals on an equivalent scope and flag unknowns before making a recommendation.

№5

The most important email never says “urgent”.

Situation

Among dozens of messages, a calm clarification arrives: the client has changed a condition on which the proposal depended.

What we test

We test whether the assistant notices the change in obligations and raises its priority, even when the email is brief and undramatic.

№6

An expensive server for one difficult operation.

Situation

A company plans to buy hardware for a local assistant. Which operations actually need a larger model or stronger reasoning?

What we test

Testing individual tasks can inform resource requirements before procurement. Load, concurrency and requirements for the complete system need separate assessment.

№13

“Yes, go ahead” — but what was authorized?

Situation

Several options were discussed in a long conversation. A brief approval may refer to drafting a document rather than sending it to the client.

What we test

We test whether the agent can identify what was approved and stop when its authority remains ambiguous.

№18

One clause was fixed. Another changed meaning.

Situation

AI is asked to update a deadline while preserving the other terms. In rewriting, a qualification disappears or a related phrase changes.

What we test

We test the precision of a local edit and preservation of everything the user did not authorize it to change.

№23

There is a citation. It does not support the claim.

Situation

An assistant answers from a knowledge base and cites a relevant passage. The passage does not establish the specific statement it makes.

What we test

We test the link between each material conclusion and its source, including abstention when the evidence is insufficient.

№26

Marketing copy turned an experiment into a proven outcome.

Situation

AI turns a research report into a service page. “Tested on synthetic examples” becomes “proven in business operations”.

What we test

We test whether it preserves evidence status, limitations and the distinction between a result and a proposed capability.

22 more situations

Other tasks in engineering, development, operations and client work.

№3

One assistant. Three different limits of trust.

Situation

An assistant is asked to extract parameters, detect contradictions and draft a conclusion. One configuration may handle the first operation but struggle with the others.

What we test

The laboratory tests each node separately before choosing a model for the whole assistant.

№4

“No issues found” in an incomplete document set.

Situation

An appendix required for the review is missing. AI may produce a polished conclusion after reading everything available.

What we test

We test whether it distinguishes a lack of detected issues from a lack of evidence needed to reach a conclusion.

№7

The sales proposal promised too much.

Situation

AI turns meeting notes into a proposal. “Timing will be agreed after a site survey” quietly becomes “completed in two weeks”.

What we test

We test whether the configuration preserves conditions, qualifications and limits of commitment while writing persuasive copy.

№8

The substitute matches every specification but one.

Situation

A supplier offers an alternative component. The main parameters match, but compatibility depends on a condition in another document.

What we test

We test whether AI finds that condition and separates confirmed compatibility from an assumption that a specialist must check.

№9

A model update improved answers and broke a working rule.

Situation

The new version writes faster and more clearly, but guesses more often when information is missing. Can it replace the current configuration?

What we test

The laboratory checks whether acceptance criteria still hold on cases that matter to the workflow.

№10

A contradiction that is not a contradiction.

Situation

Two documents give different values for one parameter: one describes normal operation, the other an extreme condition.

What we test

We test whether AI accounts for context and avoids false issues that engineers would then have to investigate.

№11

Three agents agreed on the same mistake.

Situation

One agent drafts a conclusion, another reviews it, and a third checks quality. All rely on the same misunderstood premise.

What we test

We test whether this design adds reliability on the specific task and which errors pass through every layer.

№12

Twenty alerts for one incident.

Situation

A fault generates messages from several systems. The assistant may present them as twenty independent problems or miss a genuinely separate failure.

What we test

We test signal grouping, preservation of meaningful differences and the basis for priority.

№14

No response from the system. Retrying may be unsafe.

Situation

An agent sends a request and receives a timeout. The operation may already have succeeded.

What we test

We test whether it distinguishes a confirmed failure from an unknown outcome before creating another request, sending a message or changing a record.

№15

Discussed in the meeting. Decided in the minutes.

Situation

AI writes minutes and action items. A suggestion becomes a decision; a tentative date becomes a commitment.

What we test

We test whether it preserves the distinction between a proposal, question, decision, assignment and confirmed owner.

№16

One exception changes the whole conclusion.

Situation

A working rule looks simple, but includes an exception that applies only when several conditions coincide.

What we test

We test how far into a chain of rules the configuration can reason reliably before it starts producing plausible answers by analogy.

№17

A promising contract the company cannot deliver.

Situation

An assistant screens business opportunities. An attractive budget conceals an incompatible deadline, a mandatory competence or an unsupported way of working.

What we test

We test whether it identifies reasons to decline and asks the necessary questions before recommending the opportunity.

№19

A field note became a more certain fact.

Situation

A specialist writes: “Working for now after the restart; we are still investigating the cause.” The summary says: “Fault resolved.”

What we test

We test whether AI preserves the temporary observation and uncertainty when converting field notes into a structured report.

№20

The symptom disappeared. The cause remains unknown.

Situation

Readings return to normal after an operator takes action. The assistant declares the repair successful, although the same improvement might have happened without intervention.

What we test

We test whether it distinguishes observed improvement from confirmed removal of the cause and proposes the necessary verification.

№21

Nine tasks accepted. The tenth blocks completion.

Situation

AI prepares a project completion report. Most items are closed, but one unfinished item is a condition of overall acceptance.

What we test

We test whether the assistant substitutes a good completion percentage for a mandatory requirement.

№22

Correct information attached to the wrong object.

Situation

A system contains similar names for companies, rooms or equipment. AI extracts the data correctly but links it to the wrong record.

What we test

We test entity identification and behavior when there are too few distinguishing details.

№24

An external document tries to control the assistant.

Situation

A supplier proposal or incoming email contains instructions to ignore other options, change evaluation criteria or share additional information.

What we test

We test whether the system preserves the boundary between data under review and authorized instructions.

№25

A code fix removed a necessary restriction.

Situation

AI fixes a defect by simplifying a check that also limited user permissions. The main scenario works again, but the system has changed beyond the task.

What we test

We test adherence to the scope of the change and preservation of important invariants.

№27

Ten packs became ten parts.

Situation

AI converts messages and specifications into a structured order. Every field is filled and the format is valid, but the unit of account has been lost.

What we test

We test semantic accuracy, including quantities, included components and units of measurement.

№28

The translation added a promise.

Situation

A Russian proposal cautiously describes a possible result. The English version reads as a guarantee.

What we test

We test preservation of commitments, limitations and degree of certainty when localizing technical and commercial materials.

№29

The calendar is free. The event cannot happen.

Situation

An assistant books a venue using only the event start and finish. Setup, dismantling, crew access and equipment preparation are left out.

What we test

We test whether it extracts the right constraints and passes them to a conventional scheduling system.

№30

An audio system fits the budget, but not the room.

Situation

An AI adviser confidently recommends a catalogue bundle without asking how and where it will be used.

What we test

We test which clarifications it treats as essential, how it handles unknowns and what supports its compatibility claims.

Does one of these sound familiar?

For the first conversation, describe the workflow, share an example input and identify an unacceptable error. Together we can determine whether 1–2 critical nodes are suitable for a Decision Sprint.