FILE / IT & DATA MANAGER / TSHIAMISO TRUST / PARKTOWN, JHB / V.JUL-2026

A claimant-intelligence platform for a trust that already knows what it doesn't know.

Prepared by Anthony Apollis for the Chief Executive Officer and Board of Trustees. A working proposal for how the IT & Data Manager role turns silicosis and TB compensation data into found claimants, defensible payments, and a governed system trustees can actually run.

Geo-intelligence cover showing migration routes, return tracing routes, mining hubs, labour-sending centres and executive KPI tiles
Cover visual: claimant intelligence combines KPI monitoring, migration history, return/family tracing and GIS-led outreach planning.
Project materials Download the companion Excel workbook with KPI dashboard, compensation calculator, regional model and NNA/k-NN tabs. Download Excel workbook
Candidate
Anthony Apollis
Current role
Data Engineer & Analyst, On the Spot
Reports to
Chief Executive Officer
Contact
anthony.apollis@gmail.com
00Executive Dashboard & Reader Guide

What management needs to know before reading the detail

This report is insight-led: every chart, table and map answers a management question. The reader should not have to infer the point. Each section states the data used, the finding, why it matters, and the action it supports.

414,885
Registrations
Large demand exists, but registration alone does not prove claim readiness.
Action: split registered claimants by contactability, documents and medical status.
53,945
Contact made
Contact volume is far below registration volume, showing the locating problem directly.
Action: rank uncontacted records by probability of successful tracing.
26%
Medical eligibility yield
Only 31,375 of 119,555 MCP decisions are medically eligible.
Action: track screening yield as a Board KPI, not only claim volume.
595
Certified but unpaid
Final certifications exceed claims paid; the backlog is already measurable.
Action: create a daily exception queue for certified unpaid claims.
+24/day
Backlog growth signal
TCCs grow by +36/day while payments grow by +12/day.
Action: intervene before the 90-day projection reaches about 2,755 cases.
8.9%
KZN conversion
KwaZulu-Natal has meaningful volume but weak conversion from lodgement to payment.
Action: run a targeted KZN document/medical follow-up campaign.
834 / 0
Zimbabwe lodgements / payments
A real lodgement base exists with no payment output in the public table.
Action: investigate cross-border documentation, BME and payment blockers.
0.738
ANN clustering ratio
Trust/service sites are clustered rather than randomly distributed.
Action: use spatial clustering to plan mobile outreach and midpoint service coverage.

How to read the report

What
A claimant-intelligence layer above existing claims systems.
Why
The blocker is not eligibility alone; it is finding, matching, verifying and moving claims through.
When
Start with a 90-day sequence: access controls, claimant matching, outreach automation, Board dashboard.
How
Ingest records, resolve identity, score risk/opportunity, map movement, automate follow-up, govern access.
Decision
Use evidence to prioritise regions, vendors, ICT controls, field trips and payment bottlenecks.
InsightWhy it mattersManagement decision supported
Contact made is low compared with registrationsThe Trust may know many names without knowing how to reach them.Approve claimant tracing worklists and outreach prioritisation.
Eligibility yield is only 26%Medical screening capacity must be directed toward likely eligible and document-ready cases.Add eligibility-yield and deferred-case tiles to the CEO dashboard.
Backlog is growing at a net +24 cases/dayEven a small daily mismatch becomes a Board-level service risk within one quarter.Create daily certified-unpaid exception management.
Regions behave differentlyVolume does not equal efficiency; Botswana converts well, KZN and Zimbabwe need intervention.Use region-specific playbooks instead of one national campaign.
Mineworker movement is spatially patternedMigration and return routes explain where dependants and former workers may now be found.Use the map layers to plan field trips, mobile clinics and service-site coverage.
00ADetailed Key Points & Recommendations

The proposal in plain management terms

The Claimant-Intelligence Platform addresses a specific operational problem: finding and compensating former gold mineworkers and dependants affected by silicosis and occupational TB. The issue is not only whether funds exist. The issue is whether the Trust can locate the right people, prove the right claim, protect sensitive data, and make compensation decisions that are accurate and defensible.

Core problem: many claimants are difficult to trace because their employment, medical and identity records sit across fragmented sources such as old mining archives, TEBA records, medical files and call-centre histories.
Proposed response: build a governed data layer above the existing claims system so the Trust can clean, match, enrich and prioritise claimant records without replacing the operational case-management platform.
Intelligence value: cloud warehousing, identity matching, GIS targeting and executive dashboards turn scattered records into ranked outreach lists, regional priorities and payment-bottleneck exceptions.
Operational focus: outreach should be guided by evidence from high-priority regions such as Lesotho and the Eastern Cape, weak-conversion regions such as KwaZulu-Natal, and stalled cross-border records such as Zimbabwe.
Management result: the Trust gets a defensible view of who is likely eligible, what evidence is missing, where the person or dependant may now be located, and which team action should happen next.

Recommendations for improvement

Enhanced data governance

Implement formal ownership for data quality, privacy, retention, consent, audit trails and exception handling. Old archive data should be useful, but also traceable, explainable and POPIA-compliant.

Automated outreach loops

Use SMS, WhatsApp and call-centre workflows once likely claimants are identified. Feed every failed call, wrong number, missed appointment and successful contact back into the scoring model.

Role-based access controls

Separate trustee, call-centre, medical, finance, ICT and executive access. Sensitive identity, medical and banking data should be visible only to users with a clear business need.

Continuous dashboard improvement

Treat dashboards as living management tools, not static reports. When a region lags, a BME queue stalls, or certified unpaid claims rise, the dashboard should trigger a management response.

Reader takeaway: this proposal is not a technology wish list. It is an operating model for using data and intelligence to find claimants, protect their information, target outreach, and move valid claims toward defensible payment.
01The Challenge

The Trust's hardest claim isn't the one it received. It's the one nobody filed.

Tshiamiso Trust exists to compensate ex-gold mineworkers and dependants of workers with silicosis or occupational TB. The medicine and the mandate are settled. The blocker is locating people. Many moved to rural former labour-sending areas or neighbouring countries, changed numbers, or passed away without their families knowing a claim exists.

Insight: the missing-claimant problem is an intelligence problem before it is a communications problem. Decision: build a governed tracing layer that tells field teams who to find, where, and why that person is likely eligible.

Why they go missing

Mine employment records, TEBA records, ID data and medical files sit in different systems with inconsistent name spellings, mine numbers and dates. A person who worked under three different mine numbers across two provinces doesn't look like one claimant. They look like three partial ones.

Why it's a data problem, not just an outreach problem

Outreach without matching wastes field time on people already paid, already deceased with no traced dependant, or duplicated across source systems. The Trust needs to know where to look before it decides how to reach.

02Current State

What the Trust's own public figures already show

Live counters from tshiamisotrust.com's Progress Report, pulled 15 July 2026. All 8 stages of the Trust's own claims process, not a trimmed view.

Insight: the public counters are management signals, not just website statistics. They show contact gaps, medical-yield pressure, and a certified-unpaid backlog. Decision: convert them into daily operational dashboard tiles.
53,945
Contact Made
+21 yesterday
414,885
Registrations
+127 yesterday
191,592
Appointments
+35 yesterday
162,541
Lodgements
+116 yesterday
81,903
BMEs
+6 yesterday
119,555
MCPs
+278 yesterday
28,655
Final TCCs
+36 yesterday
28,060
Claims Paid
+12 yesterday

Reading it honestly: this isn't one straight pipeline

Registrations (414,885) run far ahead of Contact Made (53,945) and Appointments (191,592) exceed Lodgements. These are 8 process counters with different denominators, not one claimant flowing cleanly through 8 gates. The MCP-vs-BME question I flagged before now has a real answer: the Trust's own medical eligibility breakdown shows 119,555 MCP decisions splitting into 31,375 medically eligible, 78,144 medically ineligible, and 9,401 deferred. MCP counts decisions, including historical-TB cases that skip a fresh BME, which is exactly why it doesn't have to track BME volume 1:1.

A found data-quality note

The homepage banner still reads "over R2.5 billion paid," while the live counter beside it reads R2.73 billion. A stale hardcoded banner next to a real-time figure. Small, but exactly the kind of inconsistency a monthly ICT/data QA pass (Section 07) is built to catch before a trustee or journalist does.

Only 26% of MCP decisions (31,375 of 119,555) come back medically eligible. Screening yield, not just volume, is a KPI worth its own dashboard tile. Certified-but-unpaid backlog today: 595 cases, and Section 06 shows it's growing, not shrinking. Full workbook: companion Excel, KPI Dashboard tab.
03Why This Trust Exists

The 500,000-strong workforce this Trust was built to compensate

The claimant-location problem in Section 01 isn't a data accident. It's the direct legacy of how the gold industry recruited and housed its labour for most of a century. Understanding that system is what makes the regional data in Section 06 make sense, not just line up.

Insight: historical labour migration explains today's claimant geography. Decision: outreach should follow recruitment and return-home patterns, not only current provincial counts.

How the labour system worked

Gold was discovered on the Witwatersrand in 1886. To staff it, the mines built two recruiting networks: the Witwatersrand Native Labour Association (WNLA, "Wenela," est. 1900) drawing from Mozambique and further afield, and the Native Recruiting Corporation (NRC) drawing from Basutoland (Lesotho), Bechuanaland (Botswana), Swaziland (Eswatini) and South Africa itself. The two merged in 1977 into TEBA. The Employment Bureau of Africa, the same archive this proposal's matching pipeline depends on.

Workers signed annual contracts and lived in single-sex hostel compounds. Over 97% of a roughly 500,000-strong workforce, cut off from family for the length of each contract, then rotating home before the next one. By the 1970s, about 70% of gold-mine labour was foreign-born. That oscillation between mine hostel and rural homestead is precisely why claimants and their descendants are hardest to find in exactly the regions that supplied the labour. The regional rankings in Section 06 aren't arbitrary. They're the recruitment map, inverted.

Why the disease burden is what it is

Autopsy studies of South African gold miners found silicosis prevalence rising from roughly 3% (Black miners) and 18% (white miners) in 1975 to 32% and 22% respectively by 2007. And in miners with around 40 years of dust exposure, silicosis was found in up to 52% at autopsy.

Tuberculosis followed the same exposure. South African gold-mine TB incidence has run at roughly 3,000–7,000 cases per 100,000 workers a year. Against a South African national average of 981 and a global average of 128 per 100,000, and more than ten times the World Health Organization's emergency threshold of 250 per 100,000. Rates peaked above 4,000 per 100,000 in 1999, compounded by HIV co-infection running near 29% in the mine workforce by 2001. Silica dust independently raises TB risk, which is exactly why the Trust compensates the two diseases under one scheme rather than two.

The settlement itself: on 26 July 2019 the High Court approved a R5 billion class-action settlement across six major mining groups (African Rainbow Minerals, Anglo American SA, AngloGold Ashanti, Gold Fields, Harmony, Sibanye-Stillwater), covering mine service between 12 March 1965 and 10 December 2019, disbursed over 12 years through the Trust. Whose name, Tshiamiso, means "to make good" in Setswana. The 2026 schedule in Section 04 is that same benefit, escalated for inflation since the original 2019 base amounts (Class 2 silicosis opened at R150,000; it is R180,059.28 now).

Sources: Nelson et al., Three Decades of Silicosis: Disease Trends at Autopsy in South African Gold Miners (Environmental Health Perspectives, via PMC); Stuckler et al. and World Bank/Results UK reporting on TB in Southern African mining; Oxford Human Rights Hub and GroundUp reporting on the 2019 Tshiamiso settlement; South African History Online and Heritage Portal on WNLA/NRC/TEBA and the compound system.

04Compensation Schedule. 1 Feb 2026 to 31 Jan 2027

Disease classes and maximum benefit

Every downstream system. Matching, dashboards, outreach prioritisation. Has to be built around this schedule, because the class a claimant is certified into determines both the payment ceiling and which data fields matter for verification.

Insight: compensation class is the financial control point. Decision: encode benefit class, work-history reductions and the 30-year exception as system rules, not spreadsheet adjustments.
ClassCertification meaningMaximum (before reductions)
Silicosis Class 1Early silicosis, lung impairment up to 10%R84,027.66
Silicosis Class 2Equivalent to 1st degree silicosisR180,059.28
Silicosis Class 3Equivalent to 2nd degree silicosisR300,098.80
Silicosis Special AwardExtraordinary, severe disease conditionsR600,197.60
Dependant Silicosis ADeceased worker, silicosis primary cause of deathR84,027.66
Dependant Silicosis BDeceased worker had a Class 2 or 3 conditionR120,039.52
TB 1st DegreeQualifying underground work + 1st degree TBR60,019.76
TB 2nd DegreeQualifying underground work + 2nd degree TBR120,039.52
Historical TBOlder certificate, usually 1965–1994R12,003.95 +
Dependant TBDeceased worker, TB primary cause of deathR120,039.52
Calculation logic
final payment = gross benefit − work-history reductions − tax (if applicable)
Reductions apply for risk work at qualifying mines outside the qualifying period, at non-qualifying gold mines, or, for TB. At non-qualifying mines such as coal or platinum. Example: Silicosis Class 2 (R180,059.28) with 3 of 10 risk-work years outside the qualifying period → 30% reduction → ≈R126,041.50 before tax.
Exception the system must encode as a rule, not a manual override: 30+ years of risk work at qualifying gold mines during qualifying periods waives the work-history reduction entirely. Full class amount is paid.
05Proposed Build

A claimant intelligence layer, sitting above the claims system

Not a replacement for the case-management system. A governed data layer that turns scattered mine, medical, call-centre and geographic records into a ranked list of who to find next, and a management view the CEO and trustees can trust.

Insight: the platform must connect identity, medical, geographic and outreach intelligence in one operating view. Decision: build the first release around ranked claimant worklists and executive exception reporting.
Ingest

Legacy & source data

Mine employment history, TEBA archives, medical exam status, call-centre logs, death registrations.

Match & clean

Identity resolution

ID number, mine number, surname variants, DOB and employer fuzzy-matched; uncertain matches flagged for human review, never auto-merged.

Locate

GIS mapping

Uncontacted claimants plotted by province, district and former labour-sending village to target field teams.

Engage

Outreach automation

SMS/WhatsApp/call-centre workflows with every attempt and failure reason logged, routed through chiefs, clinics, churches, pension pay-points.

Report

Executive dashboards

Claims registered vs. paid, bottlenecks, regional outreach yield, unresolved-document exceptions. For the CEO and Board.

Questions it must answer on demand

Who is eligible but not yet contacted? Which regions carry the highest unresolved caseload? Which claimants are stuck on missing documents or missed medical exams? Where should the next outreach trip go?

Data quality as a control, not an afterthought

Duplicate detection, missing-document alerts, mine-history validation and medical-status checks run continuously. Because a bad match here doesn't just skew a dashboard, it delays or misdirects a real payment.

Geo intelligence: migration, return tracing and outreach planning

Section 06 ranks which region carries the most unresolved caseload, using the Trust's own published regional figures. This map now does three jobs: it shows the historical movement of mineworkers from labour-sending areas to gold-mining hubs, the reverse tracing problem when sick workers or dependants return home, and the field-team outreach corridors needed to find them. Great-circle distance is still calculated from real coordinates, but the map is no longer a single-purpose distance sketch.

Loading interactive map. If it does not appear, connect to the internet or use the static fallback below.
Offline/static fallback map 448 km 452 km 166 km 263 km 154 km 58 km Johannesburg / Trust HQ Carletonville Klerksdorp Welkom Barberton Mthatha Durban Maseru Polokwane Maputo Mbabane
Historical gold-mining hub (diamond)
Labour-sending centre (circle)
Worker migration route: labour area to mine hub
Return/family tracing route: mine hub back to home area
Outreach corridor: distance-ranked field trip
Trust/service site overlay
How to read it: use the layer selector to switch purposes on/off: migration history, return tracing, outreach corridors, places, and Trust sites. Use +/− or mouse wheel to zoom, drag to pan, and switch Street/Satellite imagery to inspect terrain and settlement patterns.
Operational use: the same coordinates explain how workers moved to the mines, where they returned after illness or retirement, and where outreach teams should search for claimants or dependants now.
Method: haversine distance on real coordinates (public town/city locations). Once claimant addresses are geocoded, this same model ranks outreach trips by actual kilometres, not guesswork.
Labour-sending centreNearest gold-mining hubGreat-circle distance
Mbabane, EswatiniBarberton58 km
Maseru, LesothoWelkom (Free State Goldfields)166 km
Maputo, MozambiqueBarberton154 km
Polokwane, LimpopoBarberton263 km
Mthatha, Eastern CapeWelkom (Free State Goldfields)448 km
Durban, KwaZulu-NatalBarberton452 km

Reading it operationally: an Mbabane-based field trip sits under an hour from Barberton by straight-line distance. Cheap to bundle with a routine visit. A Mthatha or Durban trip is a multi-day exercise that needs its own budget line and batched appointments, not a same-week add-on. That's the difference between an outreach calendar built on assumption and one built on a formula that recalculates itself.

06Regional Intelligence, GIS & Machine Learning

The Trust already publishes the data this needs, nobody has run the numbers

Everything below runs on the Trust's own published Progress Report (15 July 2026). Real regional table, real correlation, real clustering, a real spatial nearest-neighbor test on 17 named Trust sites, and a real trend projection. Hover any chart or map for detail; click a table row to highlight its point.

Insight: regions differ by volume, conversion, payment outcome and spatial coverage. Decision: manage regions as different operational segments instead of one generic outreach campaign.
RegionLodgementsBMEsPaymentsConversionValue PaidCluster
South Africa (aggregate)79,76144,97612,19315.3%R1,196,531,716Trust-wide aggregate
Eastern Cape36,50318,8536,40517.5%R626,818,910Anchor (high volume)
Lesotho57,56523,27410,63318.5%R1,003,860,744Anchor (high volume)
Free State17,1958,7502,65115.4%R274,121,909Emerging (mid volume)
Mozambique11,7396,3951,83315.6%R193,813,114Emerging (mid volume)
North West9,2705,9751,39515.0%R130,612,003Emerging (mid volume)
Gauteng7,0954,16873910.4%R68,240,690Emerging (mid volume)
eSwatini6,6813,3271,42621.3%R127,310,397Emerging (mid volume)
KwaZulu-Natal6,2833,2165588.9%R58,043,007Emerging (mid volume)
Botswana5,9613,8991,97533.1%R210,348,595Small / early-stage
Mpumalanga1,39263119714.2%R16,530,635Emerging (mid volume)
Limpopo87154211913.7%R9,576,705Emerging (mid volume)
Zimbabwe8343200.0%R0Stalled, needs intervention
Western Cape7963529011.3%R8,579,205Emerging (mid volume)
Northern Cape3561973911.0%R4,008,653Emerging (mid volume)
Malawi000R0Stalled, needs intervention

Source: tshiamisotrust.com/information/progress-report/, 15 July 2026. "South Africa" is the Trust's own aggregate row, excluded from clustering and correlation below.

Real validation, not coincidence. This model independently ranks Lesotho and Eastern Cape as the two anchor regions, and the Trust's own newsroom currently carries a "historic R1 billion for Basotho ex-mineworkers" story from Maseru plus an "expanded outreach campaign for KZN ex-gold miners." That matches this model flagging KwaZulu-Natal's 8.9% conversion rate as one of the weakest at any real volume.

Correlation: what actually moves with what

Lodgements, BMEs, payments and value are almost perfectly correlated. Bigger regions have more of everything. The signal worth acting on: conversion rate barely correlates with volume. Size does not predict efficiency. Hover any cell for the exact figure.

Lodgements
BMEs
Payments
Value (R)
Conversion %
Lodgements
1.00
0.99
0.99
0.99
0.32
BMEs
0.99
1.00
0.98
0.98
0.37
Payments
0.99
0.98
1.00
1.00
0.39
Value (R)
0.99
0.98
1.00
1.00
0.40
Conversion %
0.32
0.37
0.39
0.40
1.00

Clustering: four groups, one k-means model

K-means on standardised lodgement volume, conversion rate and value-per-payment splits 15 regions into four groups. Botswana leads on efficiency at 33.1% conversion. Zimbabwe is the clearest flag: 834 real lodgements, zero payments. Hover a point, or click a table row above to find it here.

0 11,513 23,026 34,539 46,052 57,565 0 2,127 4,253 6,380 8,506 10,633 Lodgements (claims filed) Payments made
Anchor (high volume)
Emerging (mid volume)
Small / early-stage
Stalled, needs intervention
Axes: x = lodgements filed, y = payments made. Circle size = Rand value paid, bigger circle, more money moved. Colour = k-means cluster, not a fixed scale.
Hover any circle for exact figures, or click a row in the table above to jump here.

Prediction: where the backlog goes if nothing changes

Final certifications grow at +36 a day. Payments grow at only +12 a day. Drag the slider to project the backlog forward at that real, live-counter rate.

today +180 days
90
Projected certified-but-unpaid backlog: 2,755 cases

What this becomes as automation

The same techniques run on a schedule, not once. A nightly job flags any region whose conversion drops more than one standard deviation below its cluster average, which would have caught KZN and Zimbabwe automatically. The backlog trend recalculates daily and alerts the CEO if the net rate turns positive for more than a week. The cluster model re-runs monthly as new lodgements land.

Method, for the record

Python (pandas, scikit-learn, matplotlib for static exports). Pearson correlation on 5 funnel metrics. K-means (k=4, standardised features, random_state=42) on lodgements, conversion rate and value-per-payment. Linear trend projection on the homepage's own daily deltas. Full dataset: companion Excel, Regional Data (Real) tab.

Spatial intelligence: Nearest Neighbor Analysis on real sites

The Trust's Progress Report names its top lodgement, appointment and BME sites with real activity counts. Average Nearest Neighbor Analysis (ANN), the standard GIS test for whether points cluster, spread randomly, or sit uniformly, was run on all 17 of those real, named sites.

0.738
ANN ratio (R)
Below 1.0 means clustered
69.9 km
Observed mean spacing
vs 94.7 km expected at random
-2.07
z-score
Below -1.96: statistically significant clustering
17
Real sites tested
Lodgement, appointment & BME offices
Loading interactive map. If it does not appear, connect to the internet or use the static fallback below.
Offline/static fallback map Welkom Butterworth Umtata Maseru Botshabelo Molepolole Witsieshoek Mbabane Carletonville Stilfontein Mafeteng Hlotse Mohale's Hoek Teyateyaneng Maputo Orkney Gaborone
Real Trust site, circle size = activity weight
Nearest-neighbor pair (dashed line)
How to read it: use the layer selector to toggle service sites and nearest-neighbor links. Zoom in/out, drag the map, and switch to Satellite imagery to inspect the terrain around offices and labour-sending areas. The view is bounded to the real Trust-site footprint. Bigger circle = more combined appointments, lodgements and BMEs at that office. Click any site for its exact weight and nearest neighbour.
The Lesotho cluster (Maseru, Teyateyaneng, Hlotse, Mafeteng, Mohale's Hoek) sits 30-45 km apart, tightly packed. Mbabane and Maputo are 148 km apart despite both being cross-border sites, a real coverage gap in that eastern corridor worth a midpoint office.

Method: haversine pairwise distance, expected mean distance under complete spatial randomness (CSR) over the sites' bounding-box area, z-test per the standard ANN formula (as used in ArcGIS's Average Nearest Neighbor tool). Region-level version: companion Excel, NNA & k-NN (Real) tab.

Why nearest-neighbor thinking solves the missing-claimant problem

Section 01 named the real blocker: fragmented records make one person look like three. Nearest Neighbor Analysis has two forms, and both apply here.

Fragmented source data
TEBA, mine records, ID data
Spatial clustering (GIS)
Hotspot villages, mobile clinic targeting
Resolved claimant identity
Defensible payout

Spatial NNA: informal settlements and job access

Migrant workers in mining regions like North West and Free State typically settle near mine shafts in informal or shack settlements, not the formal towns in Section 06's site map above. Research on migrant labour markets shows physical proximity carries real weight: workers living closer to already-employed neighbours hear about vacancies and recruitment notice boards sooner, and get referred faster. Spatial NNA is the tool that tests this. Run against settlement GPS points (not yet in a public dataset, the gap a future field survey or drone/satellite imagery pass would close), it would prove or disprove whether tight settlement clustering near a shaft head gives job seekers a measurable time-to-employment advantage, the same clustering logic already proven significant on the Trust's own 17 real sites above.

Chain migration: predicting where a claimant worked

A migrant's most useful "nearest neighbours" are often social, not physical: kinship, home language, village of origin. Migrant mining history runs on chain migration, a worker from a specific district travels to a specific mine because a relative or fellow villager is already there. That is a k-NN problem: encode a claimant's home country, language, village and education level as features, and their nearest social neighbours in the existing employment record predict which mine or shaft they most likely worked at, even when the claimant's own mine number is missing or illegible. The same feature-distance method used for identity resolution below, pointed at a different question.

Job-seeking aspectHow NNA measures it
Physical housingChecks whether migrants cluster near mine hubs to minimise travel cost and maximise notice-board access.
Hiring pipelinesTests whether living near employed miners raises a job-seeker's chance of being referred for a vacancy.
Support networksMaps the proximity of informal or unregistered miners to host communities for resource pooling and survival.
Missing claimantsSame math, run on lodgement and payment records instead of job referrals, to find where the unexamined still live.

Fixing broken records with k-NN

Exact database joins fail on a misspelled name or a partial ID. k-NN treats each record as a point in feature space (birth year, name, mine number, region) and merges whichever records sit closest together, not just the ones that match exactly. Run on 8 synthetic demonstration records built to mimic Section 01's example, the model correctly clustered the three-partial-record case at distance 0.02 to 0.05, tight, while showing why naive tokenization still needs fuzzy name-matching in production: a shortened first name ("S" vs "Sipho") pushed one true match further out at distance 0.53. That's the argument for edit-distance matching, not simplistic hashing, in the real build.

Finding dependants through kinship proximity

Dependant Silicosis and TB claims exist because a worker died and the family never knew. "Proximity" here means shared surname, shared rural local authority, or the same TEBA recruitment batch. When a deceased, uncompensated miner's record surfaces, the same nearest-neighbor logic scans current identity and mobile registries for the closest social matches, relatives sharing lineage and geography, giving field teams an algorithmic starting point instead of a cold call to an entire village.

RecordSourceName as recordedNearest neighbork-NN distance
R1TEBA archiveSipho MokoenaR20.021
R2Mine payrollS. MokwenaR10.021
R3Call-centre intakeSipho Mokoena SnrR10.041
R4Death registrationS MokoenaR20.526
R5TEBA archiveThabo RamaboeaR60.247
R6Mine payrollT. RamabowaR50.247
R7Call-centre intakePetrus NkosiR80.33
R8Mine payrollP. NkoziR70.33

Synthetic demonstration data (8 fabricated records), not real claimants. scikit-learn NearestNeighbors, Euclidean distance on 4 encoded features. Lower distance = stronger match. Full script and output: companion Excel, NNA & k-NN (Real) tab.

07ICT Management & Governance

Five ICT problems this role inherits, and how each gets solved

Not abstractions: these map directly to the role description's ICT Management KPA and to gaps a first-round interview would surface fast.

Insight: the same data that improves compensation also increases privacy and cyber risk. Decision: implement zero trust, vendor controls, audit trails and disaster recovery as part of the data build, not after it.
ProblemHow it gets solvedHow users experience the support
Everyone works remotely across Azure/Tableau/cloud tools; a "can't access the system" call today has nowhere defined to land. Single intake channel, triaged in under 30 minutes against four standard failure classes. Credential/MFA, device, network, application. Each with a written runbook, so resolution doesn't live in one person's head. Every ticket logged start to close with a plain-language note; monthly rollup shows the CEO recurring failure types, not just a ticket count.
ICT service-provider invoices get approved without a clear check against contracted deliverables. A vendor register. Contract terms, SLA clauses, renewal dates. Every invoice reconciled against it before approval. Trustees see one vendor scorecard per quarter, not a stack of unexplained invoices.
The custom case-management system has no single business owner accountable for its data model, so changes happen ad hoc. Direct ownership, using the same audited, UAT-gated migration method already run at On the Spot. Never a big-bang cutover. Every system change ships with a short release note in plain language, not a changelog only IT can read.
Distributed, non-technical users. Trustees, call-centre staff, community liaison. Mean low adoption and repeat support questions. Role-specific one-page quick guides plus a standing weekly office-hours slot, not a single onboarding session that's forgotten by month two. The same three recurring questions get answered once, in writing, for everyone. Not re-explained on every call.
Governance committees have no regular view of ICT risk, incidents or system health. Monthly ICT dashboard. Uptime, ticket volume and resolution time, cybersecurity posture, open risks. Built on the same BI stack as the claimant dashboards in Section 05. The Board sees ICT health in the same report format it already trusts for claims data.

Access model: zero trust applied to five people who touch the same record

Zero trust means nobody is trusted by default, including staff already inside the network. Every access is verified, scoped, and logged. For the Trust, that's the difference between a call-centre agent and a trustee seeing the same claimant file.

Call-centre agent

Contact details and claim status only. No medical record, no payment approval, no bulk export.

Medical partner

Medical certification fields for assigned cases. No financial data, no other partner's caseload.

Finance user

Payment and banking data. Read-only on medical outcome. Cannot edit the class that drives the payment.

Trustee / executive

Aggregate dashboards and reports. Bulk personal data export requires a logged, approved request.

IT & Data Manager

Full system visibility for administration, under MFA, conditional access and continuous audit logging. Access itself is reviewed on a schedule, not assumed permanent.

Controls stack

MFA on every account, role-based access mapped to the table above, conditional access by device and location, encryption at rest and in transit, and audit trails on every read of medical or identity data. Because POPIA treats health and biometric data as special personal information requiring explicit safeguards.

Disaster recovery

Cloud hosting doesn't remove the need for backup discipline: automated, geographically separated backups of claimant, medical and payment data, a defined recovery-time and recovery-point objective, and a tested restore. Not just a scheduled job nobody has verified.

08Role Fit

Mapped against the five Key Performance Areas

Each KPA from the role description, with the platform mechanism that satisfies it and the tooling behind it.

Technology Governance & ICT Management
Delivered by
  • IT strategy tied to the claimant-location outcome, not technology for its own sake
  • Cloud platform reliability, security and cost ownership
  • ICT policy, standards and governance-committee input
Evidence
AzureAWSGoogle CloudData FactoryDatabricks
Business Systems & Vendor Management
Delivered by
  • Ownership of the custom claims/case-management system as business custodian
  • UAT coordination and vendor SLA/contract review discipline
  • Legacy-to-cloud migration without a service interruption to trustees
Evidence
SAP ERPSAP HANASnowflakelegacy migration
Data Management
Delivered by
  • Data governance framework and quality assurance rules on claimant records
  • Warehousing and BI strategy purpose-built for the tracing use case
  • POPIA compliance embedded in access design, not bolted on after
Evidence
Power BITableauPythonGISSnowflake ML
Strategic Planning & Team Leadership
Delivered by
  • Multi-year technology roadmap and IT/data budget ownership
  • Mentoring an IT/data team and building organisational data literacy
  • Change adoption led by user needs first. Systems fail when people aren't brought along
Evidence
stakeholder mgmtphased rollouttraining programs
Reporting & Compliance
Delivered by
  • Monthly/quarterly ICT and data reports to the CEO and Board
  • Cybersecurity posture, incident and utilisation reporting
  • Audit-ready records of contracts, licences and technology assets
Evidence
executive dashboardsaudit trails
09Track Record

Fourteen years, one throughline: messy data into decisions people can act on

2012–2014

MultiChoice Internet Holdings (Kalahari)

First full-time role after university; foundation in commercial data operations.

2014–2015

Bytes Universal Systems: BI Analyst/Developer

Business intelligence development and reporting.

2015–2016

Shoprite: SAP Master Data

SAP ERP and SAP HANA across multiple modules.

2016–2018

Macro Plan Town & Regional Planners: GIS Analyst

Mine-area planning and 3D spatial drawing. The direct precedent for claimant-location mapping.

2020–2023

TMI: Digital Marketing BI Analyst / Data Lead

Freelance analytics through the Corporate Renaissance transition and COVID period.

2023–2024

CHEP: Global logistics

Data operations inside a multinational logistics environment.

2024–Present

On the Spot: Data Engineer & Analyst

Most senior data resource in the company: legacy-to-Snowflake migration, Azure/Databricks pipelines, Power BI reporting. The same build this proposal describes, at smaller scale.

10Delivery Sequence

This is buildable, not theoretical: here's the 90-day sequence

Everything in this document. The schedule logic, the access matrix, the distance model. Is already runnable against real data. This is the order it goes into production.

Days 1–30

Access & identity foundation

MFA/RBAC live for the five roles in Section 07. First-pass ID/mine-number matching running against existing case data. First data-quality exception report to the CEO.

Days 31–60

Outreach goes live

Region-level model in Section 06 extended to claimant level: outreach automation targets actual uncontacted individuals within each region, not just the regional aggregate. Call-centre and WhatsApp outreach logged against a single claimant ID.

Days 61–90

Governance closed out

Disaster-recovery restore tested end to end. POPIA access audit signed off. First executive dashboard cycle delivered to the Board.

Anthony Apollis, Data Engineer & Analyst

anthony.apollis@gmail.com