Ingesting a market nobody sells data for
Building the data layer for a Ghanaian capital markets platform
August 2026
If you want US equity prices, you buy them. A dozen vendors will sell you a clean, versioned, documented feed with an SLA attached.
There is no equivalent for Ghana.
The Ghana Stock Exchange publishes daily trading data on its website. The Bank of Ghana publishes interbank forex rates and treasury security results on separate pages in separate formats, some as downloadable CSVs and some rendered by JavaScript. The Ghana Fixed Income Market publishes daily secondary market bond trades as Excel workbooks linked from a media library. The Ghana Statistical Service publishes CPI and GDP figures monthly and annually, inside PDF bulletins.
Four institutions, four publishing habits, no commercial feed, no contracts, and no notice when any of it changes.
Baobab, the platform I built at KAN Asset Management, needed all four in one place, valued continuously against advisor-managed client portfolios. This is how the data layer worked and what I learned building it.
Decide the approach per source, before writing the pipeline
Before writing anything permanent I spent two days on feasibility spikes. Fetch each source with a plain HTTP request and BeautifulSoup. Then with a headless browser. Then check whether an official export exists. Keep every failure artefact, the screenshots and the raw HTML and the exported files, for reference later.
That produced the single most valuable decision in the project: use the boring path wherever the source offers one.
The Bank of Ghana forex page looks dynamic, but it has a CSV export. Fetching and parsing that CSV is faster, cheaper, more stable and far less likely to break than driving a browser through the rendered page. GFIM publishes Excel workbooks, so the ingestion path is an HTTP download and a workbook parse, not scraping a table. Only where a source genuinely required JavaScript rendering did a headless browser get used at all. Reaching for the browser first is the obvious move and it is usually the wrong one. It is slower, it needs a browser inside your deployment image, it fails in more ways, and it costs meaningfully more to run on a schedule. Two days of spikes is cheap insurance against building an entire pipeline on the expensive option.
A scraper that succeeds is not the same as data that is correct
The failure mode that actually hurts is not the scraper crashing. A crash is loud, and you fix it that morning. The dangerous case is the scraper returning HTTP 200 with a page full of last week's numbers.
Two mechanisms handle this.
Freshness validation. Every forex ingestion checks the date carried by the data itself, not the date it was fetched. Data from today is current. Data inside a small tolerance window is acceptable. Anything older is flagged stale and rejected rather than written. The distinction matters because a source that quietly stops updating looks identical to one that is working, right up until somebody notices a portfolio has been valued at the wrong price for a week.
A hybrid scraper with an explicit fallback. The forex scraper tries its primary path with a timeout and a small number of retries. If every attempt fails, it activates a secondary parsing path against the same source. The resulting record is stored with a flag noting which method produced it, and the fallback activating raises an alert.
That flag is the part I would argue for hardest. Silent degradation is the enemy. If the system has been quietly running on its fallback for three weeks, you want that visible in the data itself, not buried in a log file nobody opens.
Gap detection: assume you will miss days
Any scheduled collection process will miss windows. The host restarts. The source has an outage. A public holiday breaks an assumption nobody wrote down. A deploy lands at exactly the wrong minute. Across months this is a certainty, not a risk.
The failure it creates is specific and expensive. A missing day in a price series is not a missing row. It is a wrong portfolio valuation, a wrong daily profit and loss figure, and a performance chart with a hole in it, for every client holding that instrument. Nobody notices immediately, which is exactly what makes it costly.
So the collection run does not begin by collecting. It begins by asking what is missing. A nightly scan walks backwards across a rolling window of business days and checks, per data type, whether a record exists for each date. Missing dates are cached with a short expiry and used to drive recovery. The scan is deliberately cheap, a set of count queries rather than anything analytical, because it runs ahead of every collection cycle.
Two things I only learned by running this in production:
Encode known permanent gaps. GFIM has a stretch of roughly forty weekdays in early 2023 where no reports were ever published. That is not an ingestion failure, that is the archive. Without encoding it explicitly, the gap scan flags those same forty dates as missing on every run, forever, and your alerting becomes noise you train yourself to ignore. Investigate an anomaly once, then teach the system the answer.
Tier the recovery. Recent gaps are recoverable from the daily source, so the nightly run handles them directly. Older gaps need a deeper historical crawl, which is expensive and does not belong inside a nightly job. That runs weekly, off hours, with explicit bounds on how far back it will reach and how many pages it will fetch.
Idempotency is what makes backfill safe
None of the above is usable unless re-running ingestion is free.
Two patterns cover it. Every ingested file is tracked in a per-source ingestion table keyed on its source URL with a status, so a file already ingested successfully is skipped rather than re-parsed. And every row is written as an upsert keyed on the natural key for that data type, trade date plus security type plus instrument identifier, so reprocessing a file updates existing rows instead of duplicating them.
Once that property holds, backfill stops being a dangerous operation. You can re-run any window, any number of times, without thinking hard about it. Every recovery mechanism described above depends on it, and it is considerably harder to retrofit than to design in from the start.
There is a second benefit that took me by surprise. Sources revise published figures more often than you would expect. An upsert keyed on the natural key means an upstream correction flows through on the next run, instead of leaving you with a stale number that nobody will ever think to check again.
Keep collection off the API's back
Scraping and serving requests have opposite profiles. API work is short and latency sensitive. Scraping is long, bursty and memory hungry, especially with a browser involved.
Running both on the same workers means a nightly crawl degrades response times for anyone using the product at the wrong moment. So collection runs on a dedicated queue with its own worker at a concurrency of one. That keeps it away from API traffic and stops concurrent browser instances fighting over resources. Workers restart after a fixed number of tasks, because long-running scraper processes leak.
The whole nightly cycle has a single scheduled entry point that runs the gap scan and then dispatches each source's task onto that queue. One place to look when something did not run.
Be a polite scraper
Every one of these sources is a public institution publishing data as a service. The crawler obeys robots.txt, requests one page at a time with a delay and randomised jitter, throttles itself automatically, caches responses, retries only on the status codes worth retrying, and identifies itself honestly in the user agent.
That is partly courtesy and partly self interest. The fastest way to lose access to a source you depend on is to hammer it.
The one I did not see coming
Late in the project, an upstream provider began blocking requests originating from AWS address ranges. Nothing in the application had changed. The data simply stopped arriving.
The first fix was the one I could ship within a day. The price API answers browsers from any origin, so live price synchronisation moved to the client: those requests originate from real user browsers rather than a datacentre, and prices were flowing again. That kept the product working while the underlying problem was still open, which mattered more than solving it elegantly.
But a browser relay only works while someone has the app open, and charts and snapshots need prices whether or not anyone is looking. So I spent the next three days trying to get the server its data back by engineering around the block. First a Cloudflare Worker acting as an authenticated proxy, on the assumption that the block was on AWS's address range rather than on us specifically. Cloudflare's edge addresses were refused too. Then a scheduled GitHub Actions workflow, fetching from a runner outside AWS and posting the prices to a sync endpoint. That did not hold either.
What closed it was a conversation. I contacted the provider directly, explained what we were building and how we were using the feed, and they whitelisted our server's address. Server-side fetching became the primary path again, and the browser relay stayed in place as a fallback.
Two lessons, and the second one took longer to learn. Your cloud provider's IP reputation is part of your integration surface: you do not control it, you cannot see it, and it can change underneath you without warning. And when a block is deliberate rather than incidental, a proxy or a relay from somewhere else is just a more expensive way to get refused. Sending an email was the cheaper and more durable fix, and I reached for it last rather than first.
What I would do differently
Define the ingestion contract before writing the first scraper. Each source grew its own module shape over time. A shared interface for fetch, parse, validate and upsert would have made the fourth source dramatically cheaper than the first.
Alert on absence, not only on errors. The system detects gaps well. It should also escalate when a source has produced nothing for longer than its normal publishing cadence, which is a different signal from a task failing loudly.
Version the parsers. When a source changes format, you want to know which parser version produced which historical rows. Reprocessing an archive is much harder without that, and format changes are not rare.
Ronard Adu-Botchway is a full-stack engineer in Accra, Ghana. He was the sole engineer on Baobab Capital Markets at KAN Asset Management, from first commit through to production.