Skip to content

2026-09-22 Benchmarkmaxxing

https://api-stg.opencitations.net/skg-if/v1/products?filter=cf.search.title:coviddi

Descrizioni dei servizi

Index

Meta

VoID

Ho trovato un bug di performance. RAMOSE usava una sola sessione HTTP condivisa (requests.Session) e il valore predefinito della libreria teneva aperte solo dieci connessioni. Oltre quella soglia requests butta via la connessione TCP e ne crea una nuova. Ho sovrascritto con 100.

Input: un DOI
Output: OMID dell’opera richiesta, citazioni in entrata e metadati delle opere citanti
Query remote: una federazione su un singolo articolo
Campionamento: 100 DOI, 20 con 0 opere citanti, 20 con 1-9, 20 con 10-49, 20 con 50-99 e 20 con almeno 100
Livelli di concorrenza: 1, 4, 16
Misurazioni: 6.000 = 100 DOI * 2 strategie * 10 ripetizioni * 3 livelli di concorrenza

Primi risultati, SERVICE è più veloce, con una leggera inversione oltre le 100 opere citanti. Ecco le mediane

Opere citantiConcorrenza 1, SERVICE / orchestrazioneConcorrenza 4, SERVICE / orchestrazioneConcorrenza 16, SERVICE / orchestrazione
Tutte le fasce48,5 / 65,5 ms59,1 / 80,9 ms391,1 / 398,8 ms
048,5 / 52,8 ms51,8 / 55,9 ms184,5 / 234,8 ms
1-952,4 / 57,6 ms55,0 / 63,3 ms254,3 / 293,9 ms
10-4926,2 / 69,0 ms41,9 / 77,1 ms327,3 / 355,6 ms
50-9947,0 / 86,9 ms67,0 / 101,3 ms492,7 / 501,7 ms
100+148,0 / 149,2 ms210,5 / 202,0 ms1.072,9 / 970,6 ms

E qui interviene la sofisticata arte del benchmarkmaxxing.

Stiamo mettendo sotto stress il SERVICE? Manco penniente.

Ci sono delle operazioni sulle API di OC che metterebbero sotto sforzo il SERVICE? Sì: quelle che chiedono più entità alla volta, perché il numero di chiamate al server federato aumentano.

Ok, e volendo essere ancora più stronzi? Si potrebbe fare una cosa tipo /venue-citation-count. Si parte da un ISSN, si trovano tutti gli articoli associati e per ciascuno si federa a Index

Input: un ISSN
Output: OMID degli articoli della rivista, loro citazioni in entrata e metadati delle opere citanti
Query remote: da 6 a 9.427 articoli per rivista
Campionamento: 4 ISSN scelti nelle fasce 1-9, 10-99, 100-999 e 1.000-9.999 articoli
Misurazioni: 240 = 5 ISSN * 4 fasce * 2 livelli di concorrenza * 2 strategie * 3 ripetizioni

ArticoliConcorrenzaSERVICE: latenza / successo / chiamate medie-massimeOrchestrazione: latenza / successo / chiamate medie-massime
1-91168,1 ms / 100% / 4,8-958,1 ms / 100% / 1-1
1-916304,7 ms / 100% / 4,8-9585,3 ms / 100% / 1-1
10-9913.295,4 ms / 100% / 65,6-97667,7 ms / 100% / 1-1
10-99165.028,8 ms / 100% / 65,6-974.589,8 ms / 100% / 1-1
100-999113.131,7 ms / 100% / 245,8-300794,7 ms / 100% / 1-1
100-9991615.499,1 ms / 100% / 245,8-30013.785,5 ms / 100% / 1-1
1.000-9.9991nessuna / 0% / 3.752,4-3.9831.466,1 ms / 20% / 1-1
1.000-9.99916nessuna / 0% / 3.752,4-3.98315.106,7 ms / 20% / 1-1
Già meglio. L’orchestrazione fallisce perché Virtuoso manda a QLever tutti gli OMID in un colpo solo con @@values

Aggiungo: @@values <var>... [batch_size=<positive integer>]

ArticoliConcorrenzaSERVICE: latenza / successo / richieste backend medie-massimeOrchestrazione: latenza / successo / richieste backend medie-massime
1-91200,6 ms / 100% / 5,2-987,2 ms / 100% / 3-3
1-916388,3 ms / 100% / 5,2-9439,6 ms / 100% / 3-3
10-9914.093,8 ms / 100% / 66,6-98902,5 ms / 100% / 3-3
10-99167.869,5 ms / 100% / 66,6-9810.252,4 ms / 100% / 3-3
100-999114.843,1 ms / 100% / 246,8-3011.742,0 ms / 100% / 3-3
100-9991618.219,4 ms / 100% / 246,8-30118.607,9 ms / 100% / 3-3
1.000-9.9991374.998,0 ms / 86,7% / 16.815,6-28.28477.036,5 ms / 100% / 16,6-31
1.000-9.99916153.507,0 ms / 53,3% / 16.815,6-28.284257.809,9 ms / 100% / 16,6-31
arcangelo7
arcangelo7Sep 13, 2026 · opencitations/ramose

chore(benchmark): give both stores enough memory and CPU and measure at 8 clients


Virtuoso now holds most of Meta in 44 GB of buffers and QLever runs without a CPU cap, so timeouts no longer come from a starved database. Concurrency drops from 16 to 8 because 16 heavy requests saturate the host and the benchmark would measure the machine rather than the two federation strategies.

ArticoliConcorrenzaSERVICE: latenza / successo / richieste backend medie-massimeOrchestrazione: latenza / successo / richieste backend medie-massime
1-91156,8 ms / 100% / 5,2-964,3 ms / 100% / 3-3
1-98174,3 ms / 100% / 5,2-985,7 ms / 100% / 3-3
10-9913.029,5 ms / 100% / 66,6-98118,1 ms / 100% / 3-3
10-9983.018,0 ms / 100% / 66,6-98239,9 ms / 100% / 3-3
100-999112.557,6 ms / 100% / 246,8-301215,6 ms / 100% / 3-3
100-999812.492,9 ms / 100% / 246,8-301435,6 ms / 100% / 3-3
1.000-9.9991207.908,3 ms / 100% / 5.605,2-9.42810.319,4 ms / 100% / 16,6-31
1.000-9.9998211.291,5 ms / 100% / 5.605,2-9.42813.063,5 ms / 100% / 16,6-31

E questo direi che va nell’articolo

Pasted image 20260914170916.png

Pasted image 20260914171849.png

arcangelo7
arcangelo7Sep 14, 2026 · opencitations/ramose

fix: drop the @@foreach directive


@@values covers the same ground with far fewer requests: paired with GROUP BY
it yields per-item aggregates, and it sends one query per batch instead of one
per value.

arcangelo7
arcangelo7Sep 14, 2026 · opencitations/ramose

fix: bind client values before sparql substitution


Escape literals and validate IRIs in reads, updates, and YAML templates, returning HTTP 400 for rejected placeholder values.

arcangelo7
arcangelo7Sep 18, 2026 · opencitations/ramose

feat(filters): automatically split searched values like the QLever tokenizer does if ql:has-word is found in the query

arcangelo7
arcangelo7Sep 11, 2026 · opencitations/ramose

fix(docs): preserve code casing and wrap long headings

arcangelo7
arcangelo7Sep 11, 2026 · opencitations/ramose

fix(docs): reduce horizontal margins on small screens

arcangelo7
arcangelo7Sep 11, 2026 · opencitations/ramose

fix(docs): adapt navigation to available space and align operation cards

arcangelo7
arcangelo7Sep 11, 2026 · opencitations/ramose

fix(docs): size operation cards to content

arcangelo7
arcangelo7Sep 11, 2026 · opencitations/ramose

fix(docs): up to 850px give labels and values the full available width

arcangelo7
arcangelo7Sep 11, 2026 · opencitations/ramose

fix(docs): reduce nested list indentation on small screens

arcangelo7
arcangelo7Sep 11, 2026 · opencitations/ramose

fix(docs): left-align text on small screens to avoid stretched word spacing

arcangelo7
arcangelo7Sep 11, 2026 · opencitations/ramose

feat(docs): expand JSON examples in a full-screen dialog

Pasted image 20260911201303.pngPasted image 20260911201342.png
arcangelo7
arcangelo7Sep 12, 2026 · opencitations/time-agnostic-library

feat: search provenance through a custom QLever URI index

Qvindi si può lanciare il benchmark con BEAR. Per fare un confronto, per tirare su BEAR A Fuseki impiega 34 h 39 m, QLever 1 h 55 m. BEAR A ha

Bear A non è piccolo

Dataset: 65,926,511 quad
Provenance: 1,984,865,652 quads
Totale: 2,050,792,163 quads

Ho corretto a mano le 29 br (non 42) non correggibili automaticamente

Pasted image 20260914222248.png

  • Serve una seconda cover letter per la resubmission di IEEE Access?
  • Che ne dite di switchare da Virtuoso a Qlever per Meta? L’indice testuale più o meno va. Non ci sono ancora i merge ma intanto recuperiamo un sacco di risorse (che mi servono per i benchmark di BEAR)