Repo Automation

lexmount/moli

Moli is a Rust headless browser for AI agents with on-demand layout, a CLI and CDP, WebDriver Classic and BiDi support, and low memory use.

  • 14.8k GitHub stars
  • Rust
  • ⚖️ Apache-2.0
  • 🎯 Intermediate
lexmount/moli preview image

What it is

Moli is a standalone headless browser kernel written in Rust, not a Chromium wrapper. It runs V8 JavaScript, a native DOM, CSS and networking, and builds layout and pixels only when an operation needs them. It can fetch and extract pages, search the web, and automate browser tasks through a CLI, CDP, WebDriver Classic, or WebDriver BiDi. It supports Linux, macOS, and Windows.

Who it's for

  • Developers building AI agents that need to fetch, extract, or automate web pages
  • Teams running crawling, retrieval pipelines, browser-use agents, evaluation environments, or reinforcement-learning workloads
  • Users of Playwright, Puppeteer, or Selenium who want a lightweight browser endpoint

Requirements

Requirements

  • Linux, macOS, or Windows
  • curl and a shell (Linux/macOS), or PowerShell (Windows), for the direct installer
  • Node.js with Playwright only if using the Playwright CDP example

Setup

  1. Install on Linux or macOS

    Run the installer script from the latest release.

    bash
    curl --proto '=https' --tlsv1.2 -fsSL \
      https://github.com/lexmount/moli/releases/latest/download/moli-installer.sh | sh
  2. Install on Windows

    Run the installer in PowerShell.

    powershell
    powershell -ExecutionPolicy ByPass -c "irm https://github.com/lexmount/moli/releases/latest/download/moli-installer.ps1 | iex"
  3. Install via an AI agent

    Give your AI agent a prompt that installs the Moli skills and the prebuilt binary, then fetches a page.

    text
    Install the skills under https://github.com/lexmount/moli/tree/main/skills,
    follow their instructions to download and install the latest prebuilt Moli
    binary, then use moli-webfetch to fetch https://example.com and show me the
    result.

Examples

Extract a page as Markdown

bash
bash
moli fetch \
  --dump markdown \
  --wait-until done \
  https://example.com

What it does: Renders the page as Markdown using Moli's default completion strategy.

Get a semantic tree

bash
bash
moli fetch \
  --dump semantic_tree_text \
  --wait-selector body \
  https://example.com

What it does: Returns a compact, model-friendly semantic tree once the body selector is present.

Capture screenshots and PDF

bash
bash
moli fetch --layout --dump screenshot https://example.com > page.png
moli fetch --layout --dump screenshot_full https://example.com > full-page.png
moli fetch --layout --dump pdf https://example.com > page.pdf

What it does: Adding --layout enables on-demand layout for viewport PNG, full-page PNG, and paginated PDF output.

Start the automation server

bash
bash
# Basic automation server for DOM-first workloads
moli serve

# Enable real geometry, coordinate input, and screenshot/screencast surfaces
moli serve --layout

# Also fetch optional image, font, audio, video, media, and text-track resources
moli serve --layout --resource

What it does: One endpoint serves CDP, WebDriver Classic, and WebDriver BiDi, with layout and resources enabled only when requested.

Connect Playwright over CDP

js
js
import { chromium } from "playwright";

const browser = await chromium.connectOverCDP("http://127.0.0.1:9222");
const context = browser.contexts()[0];
const page = context.pages()[0] ?? await context.newPage();

await page.goto("https://example.com");
console.log(await page.locator("body").innerText());

await browser.close();

What it does: Playwright attaches to the running Moli server at the CDP endpoint and reads the page body text.

Pros & cons

Pros

  • Pro:Structure-first design skips layout and paint unless needed; the README's sample agent workload shows 102.46 MiB peak PSS versus 348.82 MiB for Chromium
  • Pro:Single binary serves CDP, WebDriver Classic, and WebDriver BiDi, with no separate ChromeDriver, geckodriver, or browser install
  • Pro:Fetch outputs include HTML, Markdown, JSON, semantic text trees, screenshots, and PDF
  • Pro:Dual-licensed under Apache 2.0 or MIT, and usable without Lexmount Browser

Cons

  • Con:Does not pursue pixel parity with Chrome and lacks high-fidelity Canvas/WebGL/media playback
  • Con:Not every Chrome screenshot or print mode is implemented under --layout
  • Con:In the Lexbench comparison it passed 81.88% of tasks versus 99.85% for Chrome

Images