Spots

My congressional trading scraper was quietly losing half its data

I run two small Apify actors that pull U.S. congressional stock trading disclosures (the STOCK Act "PTR" filings) for the House and the Senate and turn them into clean JSON. A handful of people pay for them. Revenue had plateaued, so I did what I should have done months ago: I stopped looking at the marketing and started auditing my own output. It was worse than I thought. The number that ruined my evening I took the 50 most recent House filings and checked how many rows each one produced.

Almost half of them produced zero rows. Not

Almost half of them produced zero rows. Not errors. Not warnings anyone would see. Just nothing. The run finished green, the dataset looked healthy, and a big chunk of Congress's trades simply weren't there. After fixing it, the same 50 filings went from 238 rows to 416.

If you're a paying user, that's the kind

If you're a paying user, that's the kind of thing you'd never notice. You'd just assume your data was complete. That's what bothered me most. There wasn't one bug. There were several small ones that all failed the same way: quietly.

1. An amount format I never planned for

1. An amount format I never planned for. House filings report amounts as ranges like $1,001 - $15,000. My regex required a range. Then one filing showed up with an exact amount, $2,722.50, with cents. The regex didn't match, the parser returned an empty array, and the filing vanished.

2. Exchange transactions. Some rows are exchanges (swap

2. Exchange transactions. Some rows are exchanges (swap one holding for another), not buys or sells. I didn't have a type for them, so they got dropped.

3. Scanned paper filings. About 14% of House

3. Scanned paper filings. About 14% of House filings in the last 90 days aren't digital at all. Someone printed the form, filled it in, and hand-delivered it. The PDF is just an image. My parser saw no text, returned nothing, moved on.

4. The Senate had the same problem, and

4. The Senate had the same problem, and I had a comment saying it didn't. There was literally a code comment in my Senate actor saying "Senate has no scanned case, source is HTML." Wrong. The Senate also accepts paper filings, they just live at a different URL. In the last 30 days: 3 out of 49.

5. Network failures. When a PDF download failed

5. Network failures. When a PDF download failed after retries, I incremented an error counter and... that's it. The filing was known, it was in the index, and it disappeared from the output.

The pattern is obvious in hindsight: every one

The pattern is obvious in hindsight: every one of these paths ended in return []. And an empty array is the most polite way a program can lie to you. The fix: every filing leaves a trace

The rule I landed on is simple. Every

The rule I landed on is simple. Every filing that appears in the official index must show up in the output. Either as real transaction rows, or as a clearly marked placeholder that tells you why there's no data. So now every row has a parse_status:

News

My congressional trading scraper was quietly losing half its data

I run two small Apify actors that pull U.S.

@spots #dev
Source: Dev.to
See more like this