undefined

points

[-]

So this is a path that we definitely considered. However we think its a half-measure to generate actual Playwright code and just run that. Because if you do that, you still have a brittle test at the end of the day, and once it breaks you would need to pull in some LLM to try and adapt it anyway.

Instead of caching actual code, we cache a "plan" of specific web actions that are still described in natural language.

For example, a cached "typing" action might look like: { variant: 'type'; target: string; content: string; }

The target is a natural language description. The content is what to type. Moondream's job is simply to find the target, and then we will click into that target and type whatever content. This means it can be full vision and not rely on DOM at all, while still being very consistent. Moondream is also trivially cheap to run since it's only a 2B model. If it can't find the target or it's confidence changed significantly (using token probabilities), it's an indication that the action/plan requires adjustment, and we can dynamically swap in the planner LLM to decide how to adjust the test from there.

by ekzy287 days ago|

parent|

[-]

Did you consider also caching the coordinates returned by moondream? I understand that it is cheap, but it could be useful to detect if an element has changed position as it may be a regression

by anerli287 days ago|

parent|

[-]

So the problem is if we cache the coordinates and click blindly at the saved positions, there's no way to tell if the interface changes or if we are actually clicking the wring things (unless we try and do something hacky like listen for events on the DOM). Detecting whether elements have changed position though would definitely be feasible if re-running a test with Moondream, could compared against the coordinates of the last run.

by chrisweekly287 days ago|

parent|

[-]

sounds a lot like snapshot testing

by tomatohs287 days ago|

prev|

[-]

This is exactly our workflow, though we defined our own YAML spec [1] for reasons mentioned in previous comments.

We have multiple fallbacks to prevent flakes; The "cheap" command, a description of the intended step, and the original prompt.

If any step fails, we fall back to the next source.

1. https://docs.testdriver.ai/reference/test-steps