Skip to main content
Glama
bwalks

lcao-mcp-server

by bwalks
README.md
# lcao-mcp-server

An MCP server for searching Lackawanna County, PA real property (parcel)
records via the county Assessor's Office public website:
https://lcao.lackawannacounty.org/search/commonsearch.aspx?mode=realprop

## Why this had to be reverse engineered

That site is not an API -- it's Tyler Technologies' "iasWorld Public
Access", a classic ASP.NET WebForms app. There's no documented endpoint;
everything is server-rendered HTML with ViewState postbacks, and clicking
into a parcel's detail page depends on a server-side "search session" your
browser just created. client.py's module docstring documents the whole
flow in detail (disclaimer handshake, form fields, single-match redirect,
the sIndex/idx scheme, and the "Datalet" tab pages), reconstructed by
driving a real browser against the site and inspecting the DOM, form
fields, and navigation as each action fired.

The short version:

1. A fresh session has to POST "Agree" to a disclaimer page before the
   search form will render.
2. The Basic Search form posts back to itself with the usual WebForms
   hidden fields (__VIEWSTATE, __VIEWSTATEGENERATOR, __EVENTVALIDATION)
   plus fields for parcel #, owner name, street, and municipality.
3. A search that matches exactly one parcel redirects straight to that
   parcel's detail page. A search with multiple matches shows a results
   table instead, where every row encodes a link to its own detail page.
4. The detail page's tabs (Owner, Sales, Values, Land, etc.) are just
   plain-link navigations reusing the same session-scoped indices, and --
   very conveniently -- every data table on those pages has its HTML id
   set to the section's heading (e.g. id="Owner History"), with a
   consistent pair of CSS classes for label/value cells. That means one
   generic table parser (generic_parse_tables in client.py) can pull
   structured data out of every tab without per-tab scraping code.

Because those session-scoped indices only make sense within the browser
session that created them, this server does a fresh search-and-follow
sequence inside each tool call rather than trying to keep server-side
search state alive between MCP calls. That keeps the MCP server itself
stateless from the client's point of view, at the cost of one extra
request when you look up a parcel by its ID (cheap, since an exact parcel
number match redirects directly to the detail page in one hop).

## Installing this from GitHub

    git clone https://github.com/<your-username>/lcao-mcp-server.git
    cd lcao-mcp-server
    python3 -m venv .venv
    source .venv/bin/activate
    pip install -r requirements.txt

Then skip to "Using it from Claude Desktop / Claude Code" below, pointing
the config at wherever you cloned the repo.

## Setup (if you already have the source, e.g. via git clone above)

    cd lcao-mcp-server
    python3 -m venv .venv
    source .venv/bin/activate
    pip install -r requirements.txt

## Running it standalone (for testing)

    python server.py

This starts the MCP server on stdio. To sanity-check the scraping logic
without an MCP client in the loop, use the plain client directly:

    python -c "
    from client import LcaoClient
    with LcaoClient() as c:
        print(c.search(owner='SMITH', page_size=15))
    "

## Using it from Claude Desktop / Claude Code

Add to your MCP config (e.g. claude_desktop_config.json):

    {
      "mcpServers": {
        "lcao-property-search": {
          "command": "/absolute/path/to/lcao-mcp-server/.venv/bin/python",
          "args": ["/absolute/path/to/lcao-mcp-server/server.py"]
        }
      }
    }

## Tools exposed

- list_municipalities() -- name -> internal code for all 39
  municipalities, for the municipality argument below.
- search_by_owner(owner_name, municipality=None, page=1, page_size=50)
- search_by_address(street, municipality=None, page=1, page_size=50)
- search_by_municipality(municipality, page=1, page_size=50)
- get_parcel_detail(parcel_id) -- full record: current owner + owner/sale
  history, sales, full legal description, assessed/market values, land,
  residential characteristics, outbuildings, commercial, and notes
  (whichever sections iasWorld has data for on that parcel).

Search results carry a _sIndex/_idx pair per row -- those are the raw
session-scoped indices iasWorld uses internally. They're exposed for
debugging but aren't meaningful outside the HTTP session that produced
them, so don't try to feed them back into another tool call.

## Testing

    pip install pytest  # optional, or just run the file directly
    python tests/test_parsing.py

tests/test_parsing.py runs the HTML-parsing logic (generic_parse_tables,
form-field extraction, PARID extraction, results-list parsing, disclaimer
detection) against real markup captured live from the site, so it catches
a broken selector/regex without needing network access.

## Hardening notes

- **Street suffix normalization.** The site's address index only matches
  USPS-style abbreviated suffixes ("RD", not "ROAD") -- a spelled-out
  suffix used to come back with a hard error. `client.normalize_street`
  now rewrites common full words (Street, Avenue, Road, Drive, Boulevard,
  Lane, Court, Place, Circle, Terrace, Highway, Parkway, Trail, Square,
  Alley, Extension, Crossing, Heights, Plaza, Point, and a few more) to
  their abbreviation before submitting, so "123 Main Road" and "123
  Main Rd" both work.
- **Zero-match searches no longer error.** Observed live: a Basic Search
  with no matches doesn't show an explicit "no records" message -- it just
  redisplays the search form itself. `_interpret_search_response` now
  treats "back on frmMain with nothing else recognizable" as
  `{"type": "no_results"}` instead of raising, with a couple of literal
  no-match phrases still checked first in case the site does show one in
  some other search mode.
- **Leading directional prefixes (N/S/E/W) break the match.** Observed
  live: searching "123 N Main Rd" -- the exact way the record's own
  Property Location field is formatted -- returns no matches, while "123
  Main Rd" (directional omitted) finds it. `search()` now retries
  automatically: if a street search with a leading directional comes back
  with no_results, it strips the directional and retries once before
  giving up, and tags the result with `_note` when that's what happened.

## Known rough edges / things worth hardening before relying on this

- The parsing logic is regression-tested against real captured markup
  (see Testing above and tests/fixtures/), and every request/response
  shape in client.py's docstring was observed directly by driving a real
  browser session against the live site (form fields, redirects, tab
  navigation, pagination). What is *not* yet verified is a live end-to-end
  run of the Python client itself -- the sandbox this was built in has no
  outbound network access to the county's server, so the actual
  HTTP round trip (cookie jar behavior, and whether the real server
  accepts the exact ViewState/EventValidation replay this client does for
  paging) hasn't been executed. Run it against the live site yourself
  before relying on it; if something breaks on a first real run it's most
  likely one of those two things.
- Pagination beyond a handful of pages is slow and a little fragile. It
  replays each intermediate page as its own postback (matching how the
  site's own "Next" link works), so asking for page 20 makes 20 sequential
  requests. Fine occasionally; don't build a "dump every parcel in
  Scranton" tool on top of it without adding real rate limiting.
- No caching, retries, or backoff. This is a small county server -- keep
  request volume light and consider adding a delay/backoff if you start
  hammering it.
- get_parcel_detail makes ~10 requests per call (one per tab). That's
  inherent to how the site is built (each tab is a separate page load) but
  worth knowing if you're calling it in a loop.
- Site copyright/terms of use: the site's own disclaimer says downloaded
  material is for personal, non-commercial use (or news-media
  dissemination) and may not be resold, republished, or commercially
  exploited -- read search/commonsearch.aspx's disclaimer text yourself
  before building anything beyond personal lookups on top of this.