lcao-mcp-server
by bwalks
README.md
# lcao-mcp-server
An MCP server for searching Lackawanna County, PA real property (parcel)
records via the county Assessor's Office public website:
https://lcao.lackawannacounty.org/search/commonsearch.aspx?mode=realprop
## Why this had to be reverse engineered
That site is not an API -- it's Tyler Technologies' "iasWorld Public
Access", a classic ASP.NET WebForms app. There's no documented endpoint;
everything is server-rendered HTML with ViewState postbacks, and clicking
into a parcel's detail page depends on a server-side "search session" your
browser just created. client.py's module docstring documents the whole
flow in detail (disclaimer handshake, form fields, single-match redirect,
the sIndex/idx scheme, and the "Datalet" tab pages), reconstructed by
driving a real browser against the site and inspecting the DOM, form
fields, and navigation as each action fired.
The short version:
1. A fresh session has to POST "Agree" to a disclaimer page before the
search form will render.
2. The Basic Search form posts back to itself with the usual WebForms
hidden fields (__VIEWSTATE, __VIEWSTATEGENERATOR, __EVENTVALIDATION)
plus fields for parcel #, owner name, street, and municipality.
3. A search that matches exactly one parcel redirects straight to that
parcel's detail page. A search with multiple matches shows a results
table instead, where every row encodes a link to its own detail page.
4. The detail page's tabs (Owner, Sales, Values, Land, etc.) are just
plain-link navigations reusing the same session-scoped indices, and --
very conveniently -- every data table on those pages has its HTML id
set to the section's heading (e.g. id="Owner History"), with a
consistent pair of CSS classes for label/value cells. That means one
generic table parser (generic_parse_tables in client.py) can pull
structured data out of every tab without per-tab scraping code.
Because those session-scoped indices only make sense within the browser
session that created them, this server does a fresh search-and-follow
sequence inside each tool call rather than trying to keep server-side
search state alive between MCP calls. That keeps the MCP server itself
stateless from the client's point of view, at the cost of one extra
request when you look up a parcel by its ID (cheap, since an exact parcel
number match redirects directly to the detail page in one hop).
## Installing this from GitHub
git clone https://github.com/<your-username>/lcao-mcp-server.git
cd lcao-mcp-server
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
Then skip to "Using it from Claude Desktop / Claude Code" below, pointing
the config at wherever you cloned the repo.
## Setup (if you already have the source, e.g. via git clone above)
cd lcao-mcp-server
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
## Running it standalone (for testing)
python server.py
This starts the MCP server on stdio. To sanity-check the scraping logic
without an MCP client in the loop, use the plain client directly:
python -c "
from client import LcaoClient
with LcaoClient() as c:
print(c.search(owner='SMITH', page_size=15))
"
## Using it from Claude Desktop / Claude Code
Add to your MCP config (e.g. claude_desktop_config.json):
{
"mcpServers": {
"lcao-property-search": {
"command": "/absolute/path/to/lcao-mcp-server/.venv/bin/python",
"args": ["/absolute/path/to/lcao-mcp-server/server.py"]
}
}
}
## Tools exposed
- list_municipalities() -- name -> internal code for all 39
municipalities, for the municipality argument below.
- search_by_owner(owner_name, municipality=None, page=1, page_size=50)
- search_by_address(street, municipality=None, page=1, page_size=50)
- search_by_municipality(municipality, page=1, page_size=50)
- get_parcel_detail(parcel_id) -- full record: current owner + owner/sale
history, sales, full legal description, assessed/market values, land,
residential characteristics, outbuildings, commercial, and notes
(whichever sections iasWorld has data for on that parcel).
Search results carry a _sIndex/_idx pair per row -- those are the raw
session-scoped indices iasWorld uses internally. They're exposed for
debugging but aren't meaningful outside the HTTP session that produced
them, so don't try to feed them back into another tool call.
## Testing
pip install pytest # optional, or just run the file directly
python tests/test_parsing.py
tests/test_parsing.py runs the HTML-parsing logic (generic_parse_tables,
form-field extraction, PARID extraction, results-list parsing, disclaimer
detection) against real markup captured live from the site, so it catches
a broken selector/regex without needing network access.
## Hardening notes
- **Street suffix normalization.** The site's address index only matches
USPS-style abbreviated suffixes ("RD", not "ROAD") -- a spelled-out
suffix used to come back with a hard error. `client.normalize_street`
now rewrites common full words (Street, Avenue, Road, Drive, Boulevard,
Lane, Court, Place, Circle, Terrace, Highway, Parkway, Trail, Square,
Alley, Extension, Crossing, Heights, Plaza, Point, and a few more) to
their abbreviation before submitting, so "123 Main Road" and "123
Main Rd" both work.
- **Zero-match searches no longer error.** Observed live: a Basic Search
with no matches doesn't show an explicit "no records" message -- it just
redisplays the search form itself. `_interpret_search_response` now
treats "back on frmMain with nothing else recognizable" as
`{"type": "no_results"}` instead of raising, with a couple of literal
no-match phrases still checked first in case the site does show one in
some other search mode.
- **Leading directional prefixes (N/S/E/W) break the match.** Observed
live: searching "123 N Main Rd" -- the exact way the record's own
Property Location field is formatted -- returns no matches, while "123
Main Rd" (directional omitted) finds it. `search()` now retries
automatically: if a street search with a leading directional comes back
with no_results, it strips the directional and retries once before
giving up, and tags the result with `_note` when that's what happened.
## Known rough edges / things worth hardening before relying on this
- The parsing logic is regression-tested against real captured markup
(see Testing above and tests/fixtures/), and every request/response
shape in client.py's docstring was observed directly by driving a real
browser session against the live site (form fields, redirects, tab
navigation, pagination). What is *not* yet verified is a live end-to-end
run of the Python client itself -- the sandbox this was built in has no
outbound network access to the county's server, so the actual
HTTP round trip (cookie jar behavior, and whether the real server
accepts the exact ViewState/EventValidation replay this client does for
paging) hasn't been executed. Run it against the live site yourself
before relying on it; if something breaks on a first real run it's most
likely one of those two things.
- Pagination beyond a handful of pages is slow and a little fragile. It
replays each intermediate page as its own postback (matching how the
site's own "Next" link works), so asking for page 20 makes 20 sequential
requests. Fine occasionally; don't build a "dump every parcel in
Scranton" tool on top of it without adding real rate limiting.
- No caching, retries, or backoff. This is a small county server -- keep
request volume light and consider adding a delay/backoff if you start
hammering it.
- get_parcel_detail makes ~10 requests per call (one per tab). That's
inherent to how the site is built (each tab is a separate page load) but
worth knowing if you're calling it in a loop.
- Site copyright/terms of use: the site's own disclaimer says downloaded
material is for personal, non-commercial use (or news-media
dissemination) and may not be resold, republished, or commercially
exploited -- read search/commonsearch.aspx's disclaimer text yourself
before building anything beyond personal lookups on top of this.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues