The numbers behind the gender violence installation
— Public draft —
The gender violence installation prints a ticket every time a woman in Belgium becomes the victim of sexual violence. A name, an age, a face, a short account of what happened.
That claim only means something if the numbers behind it hold up. So this page shows the entire derivation — every step from the official statistics to the rhythm of the printer — and marks clearly which parts come straight from the research and which parts are assumed/derived otherwise.
I’m writing this up partly to get the picture straight in my own head, and because the underlying statistics are kinda slippery here and there.
The source
Everything statistical in the installation comes from a single publication:
Instituut voor de gelijkheid van vrouwen en mannen (IGVM-IEFH), Ervaringen van vrouwen en mannen met gendergerelateerd geweld in België (2026). Publication.
The raw data is the EU-GBV survey (2021), collected by Statbel, with calculations by IWEPS and UCLL. Sample: 5,494 people aged 18–74, of whom 4,529 women and 965 men (p. 19). My installation is about women only - the subsection of the report I’ll use is so women-specific that the results reported by men are statistically not relevant (doesn’t make it less horrible though).
First: what this survey actually measured
This is the part that tripped me up initially, so it’s worth explaining.
Imagine you walk into a town and ask every adult one question:
“Have you ever broken a bone?”
Say one in six raises a hand. What have you learned? That one in six people is carrying an old injury. You have not learned that one in six people break a bone this year. Most of those fractures happened five, twenty, fifty years ago.
Statisticians call the first thing prevalence — how many people carry it — and the second incidence — how many new ones happen per year. They are different numbers, and you can’t turn one into the other without knowing something extra.
The EU-GBV survey asked the “have you ever, since you were 15” question. Every headline figure in this report is prevalence. It counts scars, not fresh wounds.
The installation needs the opposite: a rate, so it knows how often to print. Getting from one to the other is the core of what follows.
The two tables everything comes from
The report is very large and touches a lot of horrible violence. However, I choose to highlight only two tables from it.
Non-partner sexual violence
Seven kinds of situations, each with the number of women who reported it. The three header rows are the totals: 471,000 women (11.5% of women aged 18–74) experienced non-partner sexual violence, of whom 220,000 rape or attempted rape and 399,000 other forms.
Note the second-to-last row — unwanted touching, 387,000. It dwarfs everything else, which is why most tickets describe groping rather than rape.
Partner sexual violence
Same structure, six kinds of situations, totals 296,000 / 282,000 / 108,000.
Why the rows don’t add up to the totals
Because a woman could tick several boxes. The seven non-partner rows sum to 779,000, but they belong to only 471,000 actual women — about 1.7 situations each. On the partner side, 602,000 across 296,000 women.
Appendix Table A.1 (p. 219) gives the real headcount: 660,000 women in Belgium aged 18–74 or 16.1% have experienced sexual violence since the age of 15.
Step 1 — Count the experiences
The installation prints events, not people. A woman who reported three different kinds of situation contributes three tickets, because three things happened to her. So the starting figure is the sum of every situation row across both tables:
non-partner 72 + 76 + 76 + 88 + 387 + 80 = 779000
partner 147 + 54 + 214 + 79 + 108 = 602000
total = 1381000 1,381,000 experiences, carried by women living in Belgium today.
That is a stock — a quantity measured at a single moment, accumulated over decades (the lifetimes of all the interviewed women). To pace a printer it has to become a flow: a quantity measured per unit of time.
Step 2 — Work out how long those took to accumulate
The tempting answer is 59 years: the survey covers ages 15 to 74, so 74 − 15 = 59.
I was wrong about this in the first iteration. 59 years is the exposure of the oldest possible respondent (74), not the average one. The survey interviewed women aged 18 to 74. A 20-year-old has been exposed to the risk for 5 years, not 59. Using 59 would treat every woman as though she were 74 — as if only the eldest had been surveyed.
We need the average exposure across all the women surveyed. The report kinda supplies it, indirectly.
Table 5.6 prints, for each age band, both a percentage and a count. Divide one by the other and you recover how many Belgian women live in that band:
18-29: 221 / 0.295 = 749,000
30-44: 230 / 0.203 = 1,133,000
45-64: 287 / 0.187 = 1,535,000
65-74: 44 / 0.065 = 677,000
total = 4,094,000 That total is a free correctness check. Table 5.3 independently states that 471,000 is 11.5% of women aged 18–74, which gives 4,096,000. Two unrelated routes, agreeing to within 0.05%. The back-calculation holds.
Now weight each band’s midpoint by its population:
(23.5 × 749) + (37 × 1,133) + (54.5 × 1,535) + (69.5 × 677)
------------------------------------------------------------
4,094
= 46.5 years The average woman in this survey is 46.5 years old, and has therefore been exposed since she was 15 for:
46.5 − 15 ≈ 31.5 years
That’s the denominator we’re looking for.
Step 3 — Divide
Experiences ÷ 31.5 years Per day
------------- -------------- -------
Non-partner 779,000 24,758/yr 68
Partner 602,000 19,133/yr 52
Total 1,381,000 43,891/yr 120 About 120 tickets a day. One every 12 minutes.
That is the installation’s real-time rhythm.
The installation uses this method of calculation but splits it in more detail according to the subcategories in the non-partner and partner tables.
We can speed up the installation to show for example one year’s worth of events in a three-month exhibition if needed.
@bert — one judgment call inside this table. Table 6.6 is headed “minstens één keer in hun leven”, not “vanaf de leeftijd van 15 jaar” like Table 5.3. Read literally, the partner stock has no age-15 floor and should be divided by mean age (46.5) rather than mean exposure (31.5), which would cut the partner rate by a third to ~12,900/yr. I applied the age-15 floor to both, since partner violence realistically starts in late adolescence (median first relationship ~17–18 gives ~29 years, close enough to 31.5). Defensible, but someone will spot it — worth being ready.
What this estimate still gets wrong
It counts kinds of experience, not incidents
This is the one that matters most, and it’s easy to state too loosely.
Each row of those tables counts the women who reported that situation at least once. So a woman groped fifty times over her life adds 1 to the unwanted-touching row, not 50. She can appear in several rows — that’s exactly why 1,381,000 is bigger than the 660,000 women who experienced any sexual violence — but repetition within a category is invisible.
So a ticket is closer to “a woman began experiencing this kind of violence” than to “an incident occurred”. Those differ by however often it repeats, which the survey mostly doesn’t ask.
Where it does ask, the gap is enormous. Among women raped by a partner, 74.1% report more than one incident within the same relationship and only 22.3% a single one. For other forms of partner sexual violence it’s 82.4% against 13.9% (Table 6.7, p. 133).
Repetition is the norm, not the exception. So measured against actual incidents, 120 a day is an undercount — and not a small one. I can’t put a number on it, because the report gives repetition counts only for partner violence and only as “once” versus “more than once”. It never says how many times.
Recall fades, and shame is a factor
People forget, and forty-year-old events get forgotten more than four-year-old ones. Sexual violence in particular carries a well-documented reluctance to disclose, which the report discusses at length in its qualitative chapters.
Both push toward under-reporting rather than over-reporting, so the true figure is more likely above my estimate than below it. But this is a general property of retrospective surveys, not something this report measures — so treat it as a direction, not a quantity.
What I can’t call a bias
It’s tempting to add a third one: women over 75, and women who have died, are outside the sample entirely, so their experiences vanish. I nearly wrote that down as a bias, and it doesn’t hold up. Their experiences are missing from the top of the fraction, but their years are equally missing from the bottom — the whole calculation is scoped to women aged 18–74 on both sides. That restricts what the rate describes; it doesn’t tilt it.
There is a residue. If violence correlates with dying earlier, then survivors under-represent victims. That’s real, probably small, and I have no way to size it from this data. Which is the honest answer, and better than dressing it up as a known bias.
And one thing that isn’t an error at all
This produces the average annual rate over the past ~31 years, not the rate in 2026. The installation smears three decades of history onto today. Rates and reporting norms have shifted across that window, so “this is happening now” really means “at the average rate of the last generation”.
That’s deliberate — it’s a fair way to render a long-running reality in the present tense — but it’s a choice, not a measurement.
None of the figures in steps 1–3 appear in the publication. They’re derived from figures that do, under the assumptions written down here. That’s a weaker claim than “the report says”, and it should be read as one.
Who is she? — Age
Each ticket carries an age. It comes from Table 5.6, the age section of which is below. The full table continues for another two-thirds of a page with education, work, health and urbanisation, which the installation doesn’t use.
Look at the Vrouwen block: two columns that disagree about which age group matters most. The % column peaks at 18–29 (29.5). The Aantal column peaks at 45–64 (287).
Both are correct; they answer different questions. The percentage says “what share of women this age are victims”. The count says “how many victims this age there actually are”.
To pick a random victim you need the counts. Here’s why, in classroom terms: in a small class of 10 children, 30% put their hand up — 3 hands. In a big class of 100, 19% put their hand up — 19 hands. If you want to pick a random hand-raiser, go to the big class. The percentage would send you to the small one.
The bands are also different widths — 45–64 spans twenty years, 65–74 only ten — which the counts already account for and the percentages don’t.
So ages are drawn in proportion to the counts:
| Age band | Share of tickets |
|---|---|
| 15–17 | 6.6% |
| 18–29 | 26.4% |
| 30–44 | 27.5% |
| 45–64 | 34.3% |
| 65–74 | 5.3% |
One ticket in three shows a woman between 45 and 64. That is what the data says, and it cuts against the reflex that victims of sexual violence are mostly young.
Two limits worth stating plainly:
The age shape is borrowed. Table 5.6 covers all non-partner violence — threats and physical violence included, not sexual violence specifically — and says nothing at all about partner violence, which is around 44% of what gets printed. The report doesn’t publish the breakdown the installation actually needs, so this is the closest available.
It’s her age now, not her age then. The survey asked women aged 18–74 about anything since they were 15. A woman assaulted at 19 and surveyed at 70 sits in the 65–74 row. So the ages describe the women carrying these experiences today, not the age at which the violence happened — consistent with the timing decision above, but it does mean the age on a ticket is not the age at which she was attacked.
The 15–17 band is mine
The survey only interviewed adults, but asked about experiences from age 15 onward. So victims aged 15–17 are present in the data and invisible in every table. I add that band at the same per-year rate as the youngest measured group, which works out to 6.6% of tickets.
@bert — this is the largest invention in the whole data layer and it needs a decision. The flat per-year rate is almost certainly too generous: a 16-year-old has had about one year of exposure since 15, while the 18–29 band averages about eight and a half. Scaling by relative exposure instead would put the band nearer 0.8% — one ticket in 125, rather than one in 15. Keeping 6.6% means minors stay visible in a piece about violence that demonstrably starts before 18; dropping to 0.8% is more defensible arithmetically. Related: the portrait prompt negatives currently exclude “child” and “teenager”, so 15–17 tickets are being rendered with adult faces either way — that mismatch needs resolving whichever number wins.
Who is she? — Name, face, origin
None of this comes from the research, and it’s the part most likely to be mistaken for data.
The report contains no breakdown of victims by origin or ethnicity. There are two passing mentions of migratieachtergrond in the literature review and nothing in any table. The installation must not be read as saying anything about which women are victims.
Faces are synthesised with Stable Diffusion via ComfyUI from a prompt template. No real person’s face appears in this work, and nobody depicted exists. Names are generated per origin, with a 50% chance of a Belgian-Dutch given name paired with a heritage surname — modelling a common second-generation pattern. Emotion, gaze, hair colour and glasses are weighted by eye, for visual variety.
@bert — the origin mix is currently described as reflecting Belgium, and it doesn’t. It runs roughly 30% non-European against a national figure closer to 17–19%, with North African at 13% versus a real ~3.5%. What’s actually been built is nearer a large-city mix than a national one. Two honest options: re-source it from Statbel origin statistics, or stop calling it Belgium-reflecting and describe it as a casting decision sampled independently of every statistic. I’d take the second — it’s what it is, and it avoids implying a link the report doesn’t support.
Was it reported?
Most tickets carry the line NIET AANGEGEVEN — not reported. It appears on 81% of them.
Two tables ask victims directly whether the police ever heard about it. They pair cleanly: both ask the same question about the most recent incident within the last five years, and both offer the same answer option — “Nee, niemand heeft het gemeld”, nobody reported it.
| Told the police herself | Nobody reported it | |
|---|---|---|
| Non-partner (Table 5.25, p. 107) | 9.7% | 84.0% |
| Partner (Table 6.24, p. 161) | 18.6% | 77.4% |
Partner violence reaches the police roughly twice as often. Since the printer emits both kinds in a fixed ratio — 779,000 non-partner experiences against 602,000 partner ones, so 56% against 44% — the two rates combine in that proportion:
(0.564 × 84.0%) + (0.436 × 77.4%) = 81.1% Two things to hold against that number. The rates describe incidents from the last five years, but they’re applied to a stock spanning about thirty, and willingness to report has shifted across that period — so 81% probably understates historical silence. And the weighting uses situation counts while the reporting tables count women, which are not quite the same unit.
It’s also a single rate standing in for two different ones, so a partner ticket and a non-partner ticket currently carry identical odds of being marked unreported when the source says they shouldn’t. An approximation, not a measurement.
What the report does say, unambiguously, is that the most common reason for not going to the police is the belief that it wasn’t serious enough or that the police weren’t the right people to tell — 58.4% for non-partner violence (Table 5.26, p. 108). For partner violence the most common reason is that it’s a private matter the woman resolved herself (43.7%, Table 6.25, p. 162).
What’s the report’s, and what’s mine
Straight from the publication:
- All thirteen situation counts, and the wording of every situation (Tables 5.3 and 6.6)
- The victim totals: 471,000 / 296,000 / 660,000
- The age-band figures (Table 5.6)
Derived from the publication, using assumptions set out above:
- The 31.5-year mean exposure window, and the population figures behind it
- 43,891 events per year — about 120 a day, one every 12 minutes
- The 81% not-reported rate
Entirely mine:
- Every story sentence. The survey collected tick-boxes, not testimony — not one sentence on any ticket is a real person’s account. They’re plausible illustrations of a category, written to make a checkbox feel like a life.
- Origin, names, faces, emotion, gaze, hair, glasses
- The 15–17 age band
- The decision to average across 31 years and print the result in the present tense
- The gaps between tickets, which are randomised rather than evenly spaced
A silence worth naming: one situation — being forced or blackmailed into sexual acts with someone else — is suppressed as VD in both tables, so it never prints. VD means “fewer than 30 women in the sample”, not “this doesn’t happen”. Appendix Table A.2 (p. 220) shows it does: 7.3% of victims, about 48,000 women. By treating it as absent, the installation inherits a silence the report only meant as a statistical precaution.
Does all of this make sense?
Corrections welcome. Every figure above traces to a numbered table in a public PDF; if I’ve got something wrong, I’d like to know.