Desktop automation

FlaUI-MCP, episode 2: driving a real application

Five lessons from driving a large WPF application through FlaUI-MCP every day: desktop sessions, low-level tools, lossy snapshots, lagging trees, real tests.

At a glance

ToolPlatformsLicenceBest for
FlaUI-MCP
  • Windows
MIT MCP server that gives AI agents Playwright-style snapshots and element refs for Windows desktop apps, built on FlaUI.
FlaUI
  • Windows
MIT .NET library over Windows UI Automation (UIA2 and UIA3); the engine underneath FlaUI-MCP.
Accessibility Insights for Windows
  • Windows
MIT Microsoft's inspector for the live UI Automation tree: properties, patterns and the raw view of any element.

Episode 1 made our fork of FlaUI-MCP reachable over HTTP, asynchronous and fast. This episode is about what happened next: driving a large WPF point-of-sale application through it every day, first with AI agents exploring screens, then with a test suite replaying what they found. Calculator demos never showed us any of the five problems below.

6. Run in an interactive desktop session, never as a service

UI Automation needs a real desktop. Windows services run in session 0, which has none, so the MCP server cannot be one — even though the reverse proxy in front of it is. It runs as an ordinary process in a logged-on session, and everything it launches opens in that session, under that account. Three consequences surprised us:

  • Screenshots come back black when the session is not rendering — disconnected, minimised RDP, or simply not the session you are looking at. Automation keeps working, because UI Automation needs no pixels, but the failure evidence attached to a report is blank.
  • A locked session refuses real input. Snapshots, focus and value fills still work, while anything that needs real input — mouse and keyboard tools, and clicks that fall back to the mouse — fails with "Access denied" ("Acceso denegado" on a Spanish Windows). The tell is that combination: reads fine, input denied, screenshots black.
  • The application runs as whoever started the server. A server left running by a colleague drives the application under the colleague's account.
On the test machine, elevated: which session holds the server, and which one is yours
qwinsta
Get-Process -IncludeUserName |
    Where-Object ProcessName -in 'FlaUI-MCP', 'FlaUI.Mcp', 'explorer' |
    Select-Object ProcessName, SessionId, UserName

No code fixes a locked desktop. Trade-off: reliable runs need test machines whose session stays logged on and unlocked, usually a dedicated account with automatic logon and no screen lock. That is a deliberate weakening of the machine's security, acceptable on an isolated lab machine and nowhere else. When a run fails with black screenshots, check the session before reading a single line of test code.

7. Add low-level tools for when the snapshot is not enough

The snapshot is designed to be short enough for a model to read: it skips unnamed decorative elements and stops at a fixed depth. That is right for most steps and wrong for some — a control visible on screen and usable by a person, yet absent from the tree the agent sees. We added six tools that go under the snapshot:

ToolWhat it does
winauto_findQuery by AutomationId, Name, ClassName or control type (contains, exact or regex), in the control or raw view, under a given ref
winauto_inspectEvery UI Automation property, pattern value and ancestor of one element; a WPF tooltip shows up as helpText
winauto_mouseHover, click, press and wheel at an element or at screen coordinates, and report the tooltip that opens
winauto_keysKey chords and navigation keys: ctrl+a, tab tab enter, alt+f4
winauto_scrollBring an item into view, set a scroll position, or turn the wheel
winauto_wait_forWait on the server until a query is present, absent, enabled or visible; the job fails on timeout

winauto_wait_for matters most. A wait the agent runs itself costs two requests per look and a model turn between them; on the server it is one job. We also added a deliberately narrow tool: the application's error dialog has a "Copy details" button that puts the full exception report on the clipboard, and that text is nowhere in the accessibility tree. The clipboard tool returns text only when it starts with that report's header, so nothing a person copied on the machine can leave it.

Trade-off: every tool lengthens the list the model chooses from and the context it carries. Keep each description precise about *when* to use the tool, not only what it does, and keep tools that touch something sensitive as narrow as the need.

8. Treat the snapshot as a lossy view, and scope every read

Two properties of real snapshots cost us more runs than anything else. First, depth is capped. Our server stops at depth 10, and anything deeper is dropped silently. Do not guess: the snapshot indents two spaces per level, so a line's depth is its leading spaces divided by two.

A grid whose checkboxes never appeared — the tell was that every cell was a leaf
Snapshot lineLeading spacesDepth
- row "SupplierModel"189
- custom "Column Display Index: 1"2010 — the cap
the checkbox inside that cell2211 — dropped

We spent four runs on "the control does not exist", "the click does not land" and "the scope is wrong" for a checkbox that was only one level too deep. The fix is to snapshot from a closer ancestor (our server added scopeRef for that), or use winauto_find. An id absent from the tree never proves the control is absent from the screen.

Second, one window can hold several screens. This application renders its dialogs as sibling groups inside a single window, so one snapshot routinely contains the local search, a cross-company search and a detail panel at once — each with its own copy of the same automation ids, Search and Cancel included. A lookup against the raw text answers for whichever copy comes first. The same happened with a "Supplier" label that also named two grid columns, and a "tick every checkbox" loop that ticked a list filter too.

  • Cut the snapshot to your screen's own subtree before matching, through one helper that every lookup goes through.
  • Anchor the cut on something unique to that screen — its title works; an automation id often does not.
  • Ask a run for the matched lines verbatim, never just a count. "2 checkboxes ticked" looked healthy while one of them was the filter.
  • WPF popups render inside the main window's HWND and never appear in the window list; splash screens are the opposite, a different HWND that the main window replaces.

9. Refs are ephemeral, and the tree lags the click

Refs are renumbered on every snapshot. Everyone knows not to reuse a ref; what is easy to miss is that the ref is part of the line. One of our retry loops recognised an element by comparing whole snapshot lines, so the line never matched, the loop concluded "gone, so it must have flipped", and it logged a checkbox as ticked that was still empty. Compare without the ref, and make "I can no longer find it" a failure, never a success:

static string WithoutRef(string line) =>
    Regex.Replace(line, @"\s*\[ref=[^\]]+\]", "");

The accessibility tree is also refreshed asynchronously. A snapshot taken right after a click can still show the old state: we had a checkbox read back as unticked while the screenshot seconds later showed it ticked. And an element can exist before its content does — a button present but still disabled, its text nodes empty. One reading is a race, not a check. Poll a few times over a short window (we use five reads over two seconds), succeed on the first read that agrees, and fail only after the last.

10. Let the agent explore, then let a test suite replay

An agent with these tools is excellent at discovery: it finds the automation ids of a screen nobody documented, works out the order of a flow, and tells you which control resists clicks. It is a poor regression suite — slower, costlier, and never exactly the same twice. So what an agent learns becomes a test: an NUnit suite that talks to the very same MCP server through the MCP C# client, one fresh MCP session per test, with page objects over the snapshot text and the polling helper from episode 1 behind every action.

Sharing one server keeps the two honest. A behaviour the agent observed is the behaviour the test gets, and a server fix benefits both. The suite then has to classify failures carefully, because only one kind deserves a retry:

FailureExampleRetry?
Infrastructure502 from the proxy, server down, job lostYes, a bounded number of times
ActionThe transport worked; the click or fill was rejectedNo — it is a finding
AssertionThe screen shows the wrong valueNo — it is the test doing its job

We learned that one the hard way: a retry attribute reran a test five times on the same machine for a rejected click, and would have marked a healthy machine unhealthy. One more habit came from the same months: before writing a test, write its scenario — data, steps, what "pass" means — and have someone who knows the business validate it. A fully coded test that tests nothing meaningful is the most expensive kind.

Where this leaves FlaUI-MCP

The snapshot-and-ref model holds up on a real application, but only once you accept that the snapshot is a summary, the tree is eventually consistent, and the desktop session is part of the test environment. If you only need an agent to drive an application on your own machine, upstream FlaUI-MCP is a few minutes from working. If you need it across machines and in a test suite, expect to make the ten changes of these two episodes — and start with the session and the transport, because every other diagnosis depends on them.

  • flaui
  • mcp
  • ai agents
  • ui automation
  • wpf
  • test design