Desktop automation

FlaUI-MCP, episode 1: taking the server remote

Five changes we made after forking FlaUI-MCP so an AI agent on a workstation can drive a Windows app on another machine: HTTP, async jobs, a faster loopback.

At a glance

ToolPlatformsLicenceBest for
FlaUI-MCP
  • Windows
MIT MCP server that gives AI agents Playwright-style snapshots and element refs for Windows desktop apps, built on FlaUI.
FlaUI
  • Windows
MIT .NET library over Windows UI Automation (UIA2 and UIA3); the engine underneath FlaUI-MCP.
MCP C# SDK
  • Windows
  • macOS
  • Linux
Apache-2.0 (MIT for older contributions) Official Model Context Protocol SDK for .NET: stdio and Streamable HTTP servers and clients from attributed tool classes.

FlaUI-MCP does for Windows desktop applications what Playwright's MCP server does for browsers. The agent asks for a snapshot of the UI Automation tree, gets back lines like button "Seven" [ref=w1e47], and clicks by ref: no screenshots to parse, no coordinates to guess. Out of the box it runs as a local stdio process next to the agent, which is perfect on one machine.

Our situation was different. The application under test — a large WPF point-of-sale application — runs on a pool of dedicated Windows test machines, and the agent runs on a developer's workstation. So we forked the project in May 2026 and spent the following months making it work across that gap, then driving a real application through it every day. This first episode covers the five server-side changes; episode 2 covers the five lessons from using it.

The starting point

Upstream builds and runs with the .NET SDK, and an MCP client launches it as a child process:

git clone https://github.com/shanselman/FlaUI-MCP.git
cd FlaUI-MCP
dotnet build src/FlaUI.Mcp
dotnet run --project src/FlaUI.Mcp

A session is a short conversation of tool calls: launch an application, take a snapshot, click or type by ref, snapshot again to check the result. Each call is a separate MCP request, and the model reads every answer before choosing the next call. Keep that in mind: it is what makes the network matter.

1. Give the tools a prefix nobody else uses

An agent sees the tools of every connected MCP server in one flat list. Upstream's tools are called windows_launch, windows_snapshot, windows_click… and "windows" is exactly the word a window-manager server or a browser server also reaches for. A collision fails loudly; a near-miss is worse, because the model quietly calls the wrong tool.

Our first commit renamed all eleven tools to winauto_*, tool descriptions and error messages included, since those are what the model reads when it recovers from a mistake. Trade-off: it breaks every existing client configuration and every prompt that names a tool. Do it on day one of a fork, not after a team depends on the names.

2. Rebuild on the official MCP SDK and .NET 10

The original server carried its own small JSON-RPC layer: a protocol class, a dispatcher and a tool registry. We replaced it with the official MCP C# SDK and moved the target framework from .NET 8 to .NET 10, the current LTS. Each tool becomes a static method with attributes, its dependencies injected as parameters, and the same tool classes serve both transports:

A tool declared through the SDK — the dependencies are injected per call
[McpServerToolType]
public static class ClickTool
{
    [McpServerTool(Name = "winauto_click")]
    [Description("Click an element by its ref.")]
    public static string Click(
        JobManager jobs,
        ElementRegistry registry,
        [Description("Element ref from winauto_snapshot")] string @ref)
    {
        var element = registry.GetElement(@ref)
            ?? throw new McpException($"Unknown ref {@ref}: re-snapshot.");
        return JobShim.Submit(jobs, "winauto_click", _ => Invoke(element));
    }
}

The deleted code was the point: the HTTP transport of the next idea came almost for free. One detail to keep from the old code — in stdio mode, stdout belongs to the protocol. Send every log line to stderr, or the first log message corrupts the JSON-RPC stream and the client disconnects with an unhelpful parse error.

3. Add Streamable HTTP, and keep it on loopback

UI Automation only works on the machine that displays the application, so the server has to run there. The question is how the agent reaches it. We weighed RDP with a local stdio server, a WinRM wrapper, raw sockets, and public tunnels, and chose the MCP standard: the Streamable HTTP transport, exposed through the reverse proxy the test machines already ran. The server binds to loopback only; the proxy terminates TLS, matches a path prefix and forwards.

One request's path
agent on the workstation
  └─ HTTPS  https://test01.example.lan:8443/flauimcp
      └─ reverse proxy on test01: TLS, strips /flauimcp
          └─ HTTP  http://localhost:5300/
              └─ FlaUI-MCP → UI Automation → application

On the agent's side there is no discovery service, just a static list of machines; the tester picks one:

VS Code .vscode/mcp.json
{
  "servers": {
    "winauto-test01": {
      "type": "http",
      "url": "https://test01.example.lan:8443/flauimcp"
    },
    "winauto-test02": {
      "type": "http",
      "url": "https://test02.example.lan:8443/flauimcp"
    }
  }
}

Check what your proxy does with streaming responses before relying on them. Ours copies a response only once the backend has finished it, so server-sent events and progress notifications never arrive incrementally. We designed for plain request and response, with progress observed by polling — which the next idea needed anyway.

4. Make every tool a job: submit, poll, cancel

The reverse proxy gives each request 30 seconds, then answers 502, and its code was not ours to change. UI Automation calls are synchronous and can block for as long as the application is busy: a snapshot of a large data grid, a launch that waits for a slow login, a batch of ten actions. So every action tool now returns a receipt immediately, and two new tools follow it up:

winauto_snapshot { "handle": "w1" }
  → Job submitted: job_51cc46e310a3bddd2a473868
    Tool: winauto_snapshot
    Poll with: winauto_job_status (job_id="job_51cc46e310a3bddd2a473868")

winauto_job_status { "job_id": "job_51cc46e310a3bddd2a473868" }
  → Status: Succeeded
    - window "Sales" [ref=w1e1] …
  • Argument checks and ref lookup still run before the job starts, so a typo fails at once instead of producing a doomed job.
  • winauto_job_cancel requests cooperative cancellation; it takes effect between UI Automation calls, never in the middle of one.
  • When an MCP session disconnects, its running jobs are cancelled — no orphaned tree walks.
  • Finished jobs stay readable for 15 minutes; job ids are 96 random bits, so one session cannot guess another's.

We made this uniform across both transports, stdio included, so there is one mental model and one test matrix. Upstream later solved the same hang differently, with a 30-second timeout per call — simpler, and enough when nothing between the agent and the server has its own timeout. Trade-off: a four-step scenario becomes 12 to 20 requests, and every client must poll. In a test suite, that polling lives in one helper:

Poll a job to its end — ModelContextProtocol client, tiered delays
static readonly Regex JobId = new(@"job_id=""([^""]+)""");
static readonly Regex State =
    new(@"^Status:\s*(\w+)", RegexOptions.Multiline);

static async Task<CallToolResult> CallAndWaitAsync(McpClient client,
    string tool, Dictionary<string, object?> args, CancellationToken ct)
{
    var receipt = await client.CallToolAsync(tool, args,
        cancellationToken: ct);
    var jobId = JobId.Match(Text(receipt)).Groups[1].Value;
    var poll = new Dictionary<string, object?> { ["job_id"] = jobId };
    for (var attempt = 0; ; attempt++)
    {
        var status = await client.CallToolAsync("winauto_job_status", poll,
            cancellationToken: ct);
        var state = State.Match(Text(status)).Groups[1].Value;
        if (state == "Succeeded") return status;
        if (state is "Failed" or "Cancelled")
            throw new InvalidOperationException($"{tool}: {Text(status)}");
        var delay = attempt < 10 ? 200 : attempt < 20 ? 2000 : 5000;
        await Task.Delay(delay, ct);
    }
}

static string Text(CallToolResult result) => string.Join("\n",
    result.Content.OfType<TextContentBlock>().Select(block => block.Text));

5. Bind to localhost, not 127.0.0.1 — and measure the transport

Test runs were green but slow, and every instinct said the UI was slow. It was not. Each request paid a flat 2.05 seconds for about 60 ms of real work, multiplied by several hundred requests per test. A ladder of three requests from the workstation finds the guilty hop without touching the UI:

PowerShell 7 — the proxy alone, one backend hop, then the MCP hop. Error statuses are expected: only the time matters
$base = 'https://test01.example.lan:8443'
$request = @{
    Method = 'Post'; Body = '{}'; ContentType = 'application/json'
    SkipHttpErrorCheck = $true
}
foreach ($path in '/nosuchroute', '/api/version', '/flauimcp/bogus') {
    $ms = (Measure-Command {
        Invoke-WebRequest "$base$path" @request
    }).TotalMilliseconds
    '{0,-16} {1,6:N0} ms' -f $path, $ms
}

We got 125 ms, 244 ms and 2,111 ms. The third request reaches a route that does not exist and executes nothing, so the jump rules out the network, the proxy, the job protocol, the snapshot size and the application in one go. The cause: the server was started with --host 127.0.0.1. Kestrel listens on IPv4 only for that literal address, but on both loopback stacks for the name localhost. The proxy forwarded to localhost, Windows tried ::1 first, and every request waited for a failed IPv6 connection before falling back.

Changing the default was not enough: a start script on several machines still passed the literal address explicitly. We now run the ladder after every server restart. Nothing had ever failed, so no correctness check would have caught it. A flat per-request cost is a timer, not work — if a trivial call and an expensive one take the same time, stop optimising the expensive one.

What it costs

ChangeWhat you gainWhat you pay
winauto_* prefixNo tool collisions between serversA breaking rename for every client and prompt
Official SDK, .NET 10Less code; HTTP from the same tools.NET 10 runtime, or self-contained builds
Streamable HTTPAgent and application on different machinesA network-exposed desktop: TLS, trusted network, auth to add
Jobs everywhereNo proxy timeout, cancellation, one model3–5× more requests; every client must poll
localhost + the ladder2 s less per requestA habit: measure after every restart

With the server reachable and fast, the hard part begins: a real application is not a calculator. Episode 2 covers what driving one every day taught us — desktop sessions, missing elements, lagging trees, and when to stop asking the agent and write a test.

  • flaui
  • mcp
  • ai agents
  • ui automation
  • dotnet