Azim Afroozeh a soutenu avec succès sa thèse de doctorat.
Daniel Lemire figure parmi les 2 % des scientifiques les plus influents au monde, selon le classement 2025 de l’Université Stanford et Elsevier. Il est également éditeur de la revue Software: Practice and Experience, une revue établie de longue date (1971) où de nombreux résultats cruciaux ont été publiés (par exemple, des articles de Knuth et Bentley). Le blogue de Daniel Lemire fait partie des 50 blogues les plus populaires sur Hacker News, le site de référence pour l’actualité technologique. Il a écrit plusieurs livres. Il siège aux comités de programme de grandes conférences en informatique.
Vous pouvez retrouver ses logiciels dans les principaux navigateurs Web, dans Git, dans les bibliothèques standard des principaux langages de programmation, et ainsi de suite. En février 2019, il a été classé en deuxième position parmi les développeurs les plus populaires sur GitHub et le plus populaire en C++ (devant Microsoft, Google et Facebook). Il fait partie des 0,0006 % de programmeurs les plus suivis au monde sur GitHub ; GitHub compte plus de 100 millions de développeurs.
En 2020 et 2021, Daniel Lemire a été coprésident du comité d’informatique du CRSNG. Il a reçu le Prix d’excellence 2020 de l’Université du Québec pour l’ensemble de ses réalisations en recherche (tous domaines confondus) pour ses travaux sur l’accélération de l’analyse JSON. Il a été admis deux fois au cercle d’excellence de l’Université du Québec.
Intérêts
- Indexation des données
- Ingénierie des données
- Performance du logiciel
- Vectorisation (SIMD)
Formation
-
Ph.D. en mathématiques de l'ingénieur, 1998
École Polytechnique et Université de Montréal
-
M.Sc. en mathématique, 1995
University of Toronto
-
B.Sc. en mathématique (mention « High Distinction »), 1994
University of Toronto (St. Michael's college)
Logiciel
« La plupart des professeurs d’informatique ne sont pas de bons programmeurs. Il existe des exceptions, comme Daniel Lemire, mais elles sont rares. » (Casey Muratori)
Je prends le développement logiciel au sérieux. On peut trouver la plupart de mes contributions logicielles sur GitHub.

Contributions choisies :
- La bibliothèque logicielle Ada est peut-être l’analyseur d’URL le plus rapide au monde. Ada a amélioré la performance du populaire environnement JavaScript Node.js : Since Node.js 18, a new URL parser dependency was added to Node.js — Ada. This addition bumped the Node.js performance when parsing URLs to a new level. Some results could reach up to an improvement of 400%. As a regular user, you may not use it directly. But if you use an HTTP server then it’s very likely to be affected by this performance improvement. (State of Node.js Performance 2023) Notre bibliothèque logicielle a aussi été adoptée par Cloudflare : It delivers a significantly faster implementation of not only URL but URLPattern and makes even more improvements on spec compliance. Elle fait aussi partie de Redpanda, de Zoom, de Telegram et d’autres systèmes importants. La bibliothèque logicielle a des fonctions avancées comme la norme URLPattern.
- simdutf : opérations Unicode plusieurs fois plus rapides que les fonctions conventionnelles.
- fast_float : lecture des nombres à virgule flottante 4 fois plus rapidement que les fonctions conventionnelles (strtod).
- simdjson : le premier analyseur JSON capable d’atteindre des vitesses de plusieurs gigaoctets par seconde, avec validation complète en utilisant un seul cœur. La bibliothèque logicielle simdjson est utilisée par Facebook, par Shopify, par Intel, par Microsoft, par Apache Doris et par plusieurs autres systèmes importants tels que Node.js. Les résultats de ces travaux sont utilisés pour accélérer le moteur HTML Blink qui équipe les navigateurs Google Chrome et Microsoft Edge, ainsi que le moteur WebKit qui équipe le navigateur Safari. La stratégie d’analyse JSON a été adoptée par le moteur JavaScript Bun chez Anthropic en 2026.
- Nos binary fuse filters sont utilisés par la plateforme X pour le filtrage rapide.
- Les bitmaps Roaring ont été largement adoptés : Google Procella (base de données de YouTube), Apache Lucene, Solr, Elasticsearch, Metamarkets’ Druid, Apache Spark, Apache Hive, Apache Tez, Apache CarbonData, Netflix Atlas, LinkedIn Pinot, Pilosa, Microsoft Visual Studio Team Services (VSTS), eBay’s Apache Kylin, et ainsi de suite. Des entreprises telles que Quantcast et Seek ont choisi les bitmaps de Roaring pour leurs besoins de performance. Lorsque Uber est passé à Apache Pinot et à ses index Roaring, ils ont économisé 2 millions de dollars par an en coûts d’infrastructure et ils ont divisé par trois le délai de chargement des pages web.
- JavaFastPFOR et FastPFor font partie de Terrier, Apache Parquet, Apache Lucene, et Apache NiFi.
- EWAHBoolArray et JavaEWAH ont été intégrés dans Git (GitHub), jGit, Apache Hive, et ainsi de suite. JavaEWAH fait partie des distributions Linux populaires comme Ubuntu et Red Hat. Les ingénieurs de GitHub ont écrit une série d’articles sur leur application des bitmaps EWAH afin d’accélérer le traitement du code. La documentation de Git traite du format EWAH.
Certains des billets de mon blogue ont mené à des améliorations au sein de logiciels bien connus.
- Mon billet Accelerating PHP hashing by “unoptimizing” it a mené à une optimisation de la fonction de hachage au sein du langage PHP (PHP 7.4).
- Mon billet A fast alternative to the modulo reduction décrit une technique utilisée par TensorFlow, par Facebook RocksDB, par Google netstack et par le noyau Bitcoin.
- Mon billet Computing the number of digits of an integer even faster a aidé à l’optimisation d’Oracle TruffleRuby et Microsoft .NET.
- Mon billet Visiting all values in an array exactly once in random order est cité dans le code source du compilateur Swift.
- Des techniques pour mesurer le parallélisme de la mémoire développées pour mon blogue ont été adoptées par des journalistes techno tels que ceux du site anandtech.com.
- L’équipe du GitHub search considère mes travaux comme étant le fondement de leur moteur de recherche.
- En 2025, nous avons multiplié par trois ou quatre la vitesse d’encodage base64 dans la célèbre bibliothèque OpenSSL pour les processeurs compatibles AVX2.
Plusieurs de nos articles scientifiques ont aussi eu un effet notable.
- Notre lecteur de nombres à virgule flottante décrit dans l’article Number Parsing at a Gigabyte per Second a été adopté par les langages de programmation C#, Go et Rust, par Apache Arrow, par Yandex ClickHouse, par Microsoft LightGBM, par le parseur JSON Jackson, et par plusieurs autres systèmes importants. Les notes de la version Go 1.16 nous informent que “ParseFloat now uses the Eisel-Lemire algorithm, improving performance by up to a factor of 2. This can also speed up decoding textual formats like encoding/json.” Les notes de la version Rust 1.55 nous informent que la “standard library’s implementation of float parsing has been updated to use the Eisel-Lemire algorithm, which brings both speed improvements and improved correctness”. L’algorithme fait partie de la bibliothèque standard C au sein de LLVM. Notre approche a été adoptée par C# à compter de .NET7. Elle fait partie de la bibliothèque C++ standard sous Linux à compter de GCC 12. Elle fait partie de la bibliothèque Mojo standard. La bibliothèque Abseil de Google a aussi adopté notre approche. Elle fait aussi partie de WebKit, le moteur de Safari, le navigateur web d’Apple. Elle a aussi été adoptée par Chromium, le moteur derrière Google Chrome et Microsoft Edge. Le moteur de base de données en mémoire Redis a également adopté notre approche qui a permis d’améliorer le temps de latence de 30% dans certains cas. Elle fait aussi partie de MySQL, Boost JSON, Blender, etc.
- Notre algorithme de compression StreamVByte décrit dans l’article Stream VByte: Faster Byte-Oriented Integer Compression est utilisé par Facebook Thrift, StarRocks et RedisLabs’ RediSearch.
- L’algorithme de génération de nombres aléatoires décrit dans mon article Fast Random Integer Generation in an Interval a été adopté
- par la bibliothèque C++ sous Linux (GNU libstdc++) pour accélérer la fonction std::uniform_int_distribution (à compter de GNU GCC 11),
- par la bibliothèque C++ de Microsoft,
- par la bibliothèque Apache Commons,
- par le noyau Linux,
- par la bibliothèque C de FreeBSD,
- par les Google’s Abseil C++ Common Libraries,
- par le langage Swift,
- par le langage Go,
- par le langage Julia,
- par le langage Zig,
- par Numpy (Python).
- L’algorithme décrit dans notre article Faster Base64 Encoding and Decoding using AVX2 Instructions est utilisé au sein du langage PHP (à compter de la version 7.4) et de la bibliothèque standard du C#. L’algorithme décrit dans notre article Base64 encoding and decoding at almost the speed of a memory copy est utilisé au sein de l’OpenJDK (Java) afin d’accélérer java/util/Base64. L’algorithme est aussi utilisé au sein de la bibliothèque standard du C#, au sein de la bibliothèque standard de Mojo et au sein du navigateur Safari.
- L’algorithme décrit dans notre article Faster Population Counts using AVX2 Instructions est utilisé au sein du Windows Terminal de Microsoft.
- L’algorithme décrit dans notre article Faster Remainder by Direct Computation: Applications to Compilers and Software Libraries est
- utilisé au sein de la bibliothèque standard du C# pour accélérer notamment la classe
Dictionaryet pour accélérer les appels de fonctions virtuelles. - Notre approche originale est utilisée par le langage Go: son application a accéléré plusieurs programmes écrits en Go par environ 1.5%.
- Elle est aussi utilisée au sein de la mise en œuvre de l’unordered map au sein de Boost contribuant ainsi à la performance remarquable.
- L’algorithme décrit dans notre article Validating UTF-8 In Less Than One Instruction Per Byte est utilisé au sein de l’interpréteur PHP, d’Oracle GraalVM et de Google Fuchsia et de plusieurs autres systèmes. Notre bibliothèque C++ simdutf, qui contient cet algorithme ainsi que de nombreux autres, fait partie de systèmes majeurs tels que Node.js, Bun et WebKit (le moteur de Safari). L’adoption de la bibliothèque simdutf par Node.js a mené à une amélioration considérable de la performance : Decoding and Encoding becomes considerably faster than in Node.js 18. With the addition of simdutf for UTF-8 parsing the observed benchmark, results improved by 364% (an extremely impressive leap) when decoding in comparison to Node.js 16. (State of Node.js Performance 2023) Il fait également partie du langage de programmation Mojo où il a rendu la validation UTF-8 dix fois plus rapide.
- L’équipe du moteur de recherche de GitHub a identifié mon travail comme étant fondamental.
- L’article On-demand JSON : A better way to parse documents? a été l’article le plus lu des 5 dernières années chez Software: Practice and Experience (2024).
- L’algorithme de l’article Batched Ranged Random Integer Generation a été adopté par la bibliothèque Apache Commons RNG. Il est aussi à l’étude pour la bibliothèque standard C++ de Microsoft, où il accélérerait par un facteur de 5 la fonction
std::shuffle.
Publications récentes
Vous pouvez trouver mes travaux sur arXiv, sur Google Scholar, sur DBLP, sur le portail ACM, sur R Libre et ailleurs.
-
Robert Clausecker, Daniel LemireFixing ill-formed UTF-16 strings with SIMD instructionsSoftware: Practice and Experience, 2026
Renseignements — Fixing ill-formed UTF-16 strings with SIMD instructions PDF (arXiv) — Fixing ill-formed UTF-16 strings with SIMD instructions DOI — Fixing ill-formed UTF-16 strings with SIMD instructions
-
Jaël Champagne Gareau, Daniel LemireConverting an Integer to a Decimal String in Under Two NanosecondsSoftware: Practice and Experience 56 (8), 2026
Renseignements — Converting an Integer to a Decimal String in Under Two Nanoseconds PDF (arXiv) — Converting an Integer to a Decimal String in Under Two Nanoseconds Code — Converting an Integer to a Decimal String in Under Two Nanoseconds
-
Jaël Champagne Gareau, Daniel LemireConverting Binary Floating-Point Numbers to Shortest Decimal Strings: An Experimental ReviewSoftware: Practice and Experience 56 (4), 2026
Renseignements — Converting Binary Floating-Point Numbers to Shortest Decimal Strings: An Experimental Review PDF — Converting Binary Floating-Point Numbers to Shortest Decimal Strings: An Experimental Review Code — Converting Binary Floating-Point Numbers to Shortest Decimal Strings: An Experimental Review
-
Robert Clausecker, Daniel Lemire, Florian SchintkeFaster Positional-Population Counts for AVX2, AVX-512, and ASIMDConcurrency and Computation: Practice and Experience 37 (27-28), 2025
Renseignements — Faster Positional-Population Counts for AVX2, AVX-512, and ASIMD PDF (arXiv) — Faster Positional-Population Counts for AVX2, AVX-512, and ASIMD Code — Faster Positional-Population Counts for AVX2, AVX-512, and ASIMD
-
Daniel LemireScanning HTML at Tens of Gigabytes per Second on ARM ProcessorsSoftware: Practice and Experience 55 (7), 2025
Renseignements — Scanning HTML at Tens of Gigabytes per Second on ARM Processors PDF (arXiv) — Scanning HTML at Tens of Gigabytes per Second on ARM Processors Code — Scanning HTML at Tens of Gigabytes per Second on ARM Processors
-
Jeroen Koekkoek, Daniel LemireParsing Millions of DNS Records per SecondSoftware: Practice and Experience 55 (4), 2025
Renseignements — Parsing Millions of DNS Records per Second PDF (arXiv) — Parsing Millions of DNS Records per Second Code — Parsing Millions of DNS Records per Second
-
Nevin Brackett-Rozinsky, Daniel LemireBatched Ranged Random Integer GenerationSoftware: Practice and Experience 55 (1), 2025
Renseignements — Batched Ranged Random Integer Generation PDF (arXiv) — Batched Ranged Random Integer Generation Code — Batched Ranged Random Integer Generation
-
John Keiser, Daniel LemireOn-Demand JSON: A Better Way to Parse Documents?Software: Practice and Experience 54 (6), 2024
Renseignements — On-Demand JSON: A Better Way to Parse Documents? PDF (arXiv) — On-Demand JSON: A Better Way to Parse Documents? Code — On-Demand JSON: A Better Way to Parse Documents?
-
Yagiz Nizipli, Daniel LemireParsing Millions of URLs per SecondSoftware: Practice and Experience 54 (5), 2024
Renseignements — Parsing Millions of URLs per Second PDF (arXiv) — Parsing Millions of URLs per Second Code — Parsing Millions of URLs per Second
-
Daniel LemireExact Short Products From Truncated MultipliersComputer Journal 67 (4), 2024
Renseignements — Exact Short Products From Truncated Multipliers PDF (arXiv) — Exact Short Products From Truncated Multipliers Code — Exact Short Products From Truncated Multipliers
-
Robert Clausecker, Daniel LemireTranscoding Unicode Characters with AVX-512 InstructionsSoftware: Practice and Experience 53 (12), 2023.
Renseignements — Transcoding Unicode Characters with AVX-512 Instructions PDF (arXiv) — Transcoding Unicode Characters with AVX-512 Instructions Code — Transcoding Unicode Characters with AVX-512 Instructions
-
Noble Mushtak, Daniel LemireFast Number Parsing Without FallbackSoftware: Practice and Experience 53 (7), 2023
Renseignements — Fast Number Parsing Without Fallback PDF (arXiv) — Fast Number Parsing Without Fallback
-
Thomas Mueller Graf, Daniel LemireBinary Fuse Filters: Fast and Smaller Than Xor FiltersJournal of Experimental Algorithmics 27, 2022
Renseignements — Binary Fuse Filters: Fast and Smaller Than Xor Filters PDF (arXiv) — Binary Fuse Filters: Fast and Smaller Than Xor Filters Code — Binary Fuse Filters: Fast and Smaller Than Xor Filters
-
Daniel Lemire, Wojciech MułaTranscoding Billions of Unicode Characters per Second with SIMD InstructionsSoftware: Practice and Experience 52 (2), 2022
Renseignements — Transcoding Billions of Unicode Characters per Second with SIMD Instructions PDF (arXiv) — Transcoding Billions of Unicode Characters per Second with SIMD Instructions Code — Transcoding Billions of Unicode Characters per Second with SIMD Instructions
-
Daniel LemireUnicode at Gigabytes per SecondSPIRE 2021: String Processing and Information Retrieval
Renseignements — Unicode at Gigabytes per Second PDF (arXiv) — Unicode at Gigabytes per Second Code — Unicode at Gigabytes per Second
-
Daniel Lemire, Colin Bartlett, Owen KaserInteger Division by Constants: Optimal BoundsHeliyon 7 (6), 2021
Renseignements — Integer Division by Constants: Optimal Bounds PDF (arXiv) — Integer Division by Constants: Optimal Bounds
-
Marcus D. R. Klarqvist, Wojciech Muła, Daniel LemireEfficient Computation of Positional Population Counts Using SIMD InstructionsConcurrency and Computation: Practice and Experience 33 (17), 2021
Renseignements — Efficient Computation of Positional Population Counts Using SIMD Instructions PDF (arXiv) — Efficient Computation of Positional Population Counts Using SIMD Instructions Code — Efficient Computation of Positional Population Counts Using SIMD Instructions
-
Daniel LemireNumber Parsing at a Gigabyte per SecondSoftware: Practice and Experience 51 (8), 2021
Renseignements — Number Parsing at a Gigabyte per Second PDF (arXiv) — Number Parsing at a Gigabyte per Second Code — Number Parsing at a Gigabyte per Second
-
John Keiser, Daniel LemireValidating UTF-8 In Less Than One Instruction Per ByteSoftware: Practice and Experience 51 (5), 2021
Renseignements — Validating UTF-8 In Less Than One Instruction Per Byte PDF (arXiv) — Validating UTF-8 In Less Than One Instruction Per Byte Code — Validating UTF-8 In Less Than One Instruction Per Byte
-
Thomas Mueller Graf, Daniel LemireXor Filters: Faster and Smaller Than Bloom and Cuckoo FiltersJournal of Experimental Algorithmics 25 (1), 2020
Renseignements — Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters PDF (arXiv) — Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters Code — Xor Filters: Faster and Smaller Than Bloom and Cuckoo Filters
-
Wojciech Muła, Daniel LemireBase64 encoding and decoding at almost the speed of a memory copySoftware: Practice and Experience 50 (2), 2020
Renseignements — Base64 encoding and decoding at almost the speed of a memory copy PDF (arXiv) — Base64 encoding and decoding at almost the speed of a memory copy Code — Base64 encoding and decoding at almost the speed of a memory copy
-
Geoff Langdale, Daniel LemireParsing Gigabytes of JSON per SecondVLDB Journal 28 (6), 2019
Renseignements — Parsing Gigabytes of JSON per Second PDF (arXiv) — Parsing Gigabytes of JSON per Second Code — Parsing Gigabytes of JSON per Second
-
Daniel LemireFast Random Integer Generation in an IntervalACM Transactions on Modeling and Computer Simulation 29 (1), 2019
Renseignements — Fast Random Integer Generation in an Interval PDF (arXiv) — Fast Random Integer Generation in an Interval Vidéo — Fast Random Integer Generation in an Interval Code — Fast Random Integer Generation in an Interval
-
Daniel Lemire, Owen Kaser, Nathan KurzFaster Remainder by Direct Computation: Applications to Compilers and Software LibrariesSoftware: Practice and Experience 49 (6), 2019
Renseignements — Faster Remainder by Direct Computation: Applications to Compilers and Software Libraries PDF (arXiv) — Faster Remainder by Direct Computation: Applications to Compilers and Software Libraries Code — Faster Remainder by Direct Computation: Applications to Compilers and Software Libraries
-
Daniel Lemire, Melissa E. O’NeillXorshift1024*, Xorshift1024+, Xorshift128+ and Xoroshiro128+ Fail Statistical Tests for LinearityComputational and Applied Mathematics 350, 2019
Renseignements — Xorshift1024*, Xorshift1024+, Xorshift128+ and Xoroshiro128+ Fail Statistical Tests for Linearity PDF (arXiv) — Xorshift1024*, Xorshift1024+, Xorshift128+ and Xoroshiro128+ Fail Statistical Tests for Linearity Code — Xorshift1024*, Xorshift1024+, Xorshift128+ and Xoroshiro128+ Fail Statistical Tests for Linearity
-
Wojciech Muła, Daniel LemireFaster Base64 Encoding and Decoding using AVX2 InstructionsACM Transactions on the Web 12 (3), 2018
Renseignements — Faster Base64 Encoding and Decoding using AVX2 Instructions PDF (arXiv) — Faster Base64 Encoding and Decoding using AVX2 Instructions Code — Faster Base64 Encoding and Decoding using AVX2 Instructions
-
Daniel Lemire, Owen Kaser, Nathan Kurz, Luca Deri, Chris O’Hara, François Saint-Jacques, Gregory Ssi-Yan-KaiRoaring Bitmaps: Implementation of an Optimized Software LibrarySoftware: Practice and Experience 48 (4), 2018
Renseignements — Roaring Bitmaps: Implementation of an Optimized Software Library PDF (arXiv) — Roaring Bitmaps: Implementation of an Optimized Software Library Code — Roaring Bitmaps: Implementation of an Optimized Software Library
-
Edmon Begoli, Jesús Camacho Rodríguez, Julian Hyde, Michael J. Mior, Daniel LemireApache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data SourcesSIGMOD'18, 2018
Renseignements — Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources PDF (arXiv) — Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources Code — Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources
-
Wojciech Muła, Nathan Kurz, Daniel LemireFaster Population Counts Using AVX2 InstructionsComputer Journal 61 (1), 2018
Renseignements — Faster Population Counts Using AVX2 Instructions PDF (arXiv) — Faster Population Counts Using AVX2 Instructions Code — Faster Population Counts Using AVX2 Instructions
-
Antonio Badia, Daniel LemireOn Desirable Semantics of Functional Dependencies over Databases with Incomplete InformationFundamenta Informaticae 158 (4), 2018
Renseignements — On Desirable Semantics of Functional Dependencies over Databases with Incomplete Information PDF (arXiv) — On Desirable Semantics of Functional Dependencies over Databases with Incomplete Information
-
Daniel Lemire, Nathan Kurz, Christoph RuppStream VByte: Faster Byte-Oriented Integer CompressionInformation Processing Letters 130, 2018
Renseignements — Stream VByte: Faster Byte-Oriented Integer Compression PDF (arXiv) — Stream VByte: Faster Byte-Oriented Integer Compression Code — Stream VByte: Faster Byte-Oriented Integer Compression
-
Daniel Lemire, Christoph RuppEfficient Integer-Key Compression in a Key-Value Store using SIMD InstructionsInformation Systems 66, 2017
Renseignements — Efficient Integer-Key Compression in a Key-Value Store using SIMD Instructions PDF (arXiv) — Efficient Integer-Key Compression in a Key-Value Store using SIMD Instructions Code — Efficient Integer-Key Compression in a Key-Value Store using SIMD Instructions
-
Dmytro Ivanchykhin, Sergey Ignatchenko, Daniel LemireRegular and almost universal hashing: an efficient implementationSoftware: Practice and Experience 47 (10), 2017
Renseignements — Regular and almost universal hashing: an efficient implementation PDF (arXiv) — Regular and almost universal hashing: an efficient implementation Code — Regular and almost universal hashing: an efficient implementation
-
Samy Chambi, Daniel Lemire, Owen Kaser, Robert GodinBetter bitmap performance with Roaring bitmapsSoftware: Practice and Experience 46 (5), 2016
Renseignements — Better bitmap performance with Roaring bitmaps PDF (arXiv) — Better bitmap performance with Roaring bitmaps Diapositives — Better bitmap performance with Roaring bitmaps Code — Better bitmap performance with Roaring bitmaps Projet — Better bitmap performance with Roaring bitmaps
-
Owen Kaser, Daniel LemireCompressed bitmap indexes: beyond unions and intersectionsSoftware: Practice and Experience 46 (2), 2016
Renseignements — Compressed bitmap indexes: beyond unions and intersections PDF (arXiv) — Compressed bitmap indexes: beyond unions and intersections Code — Compressed bitmap indexes: beyond unions and intersections
-
Daniel Lemire, Gregory Ssi-Yan-Kai, Owen KaserConsistently faster and smaller compressed bitmaps with RoaringSoftware: Practice and Experience 46 (11), 2016
Renseignements — Consistently faster and smaller compressed bitmaps with Roaring PDF (arXiv) — Consistently faster and smaller compressed bitmaps with Roaring Diapositives — Consistently faster and smaller compressed bitmaps with Roaring Code — Consistently faster and smaller compressed bitmaps with Roaring Projet — Consistently faster and smaller compressed bitmaps with Roaring
-
Daniel Lemire, Owen KaserFaster 64-bit universal hashing using carry-less multiplicationsJournal of Cryptographic Engineering 6(3), 2016
Renseignements — Faster 64-bit universal hashing using carry-less multiplications PDF (arXiv) — Faster 64-bit universal hashing using carry-less multiplications Code — Faster 64-bit universal hashing using carry-less multiplications
-
Daniel Lemire, Leonid Boytsov, Nathan KurzSIMD Compression and the Intersection of Sorted IntegersSoftware: Practice and Experience 46 (6), 2016
Renseignements — SIMD Compression and the Intersection of Sorted Integers PDF (arXiv) — SIMD Compression and the Intersection of Sorted Integers Diapositives — SIMD Compression and the Intersection of Sorted Integers Code — SIMD Compression and the Intersection of Sorted Integers
-
Wayne Xin Zhao, Xudong Zhang, Daniel Lemire, Dongdong Shan, Jian-Yun Nie, Hongfei Yan, Ji-Rong WenA General SIMD-based Approach to Accelerating Compression AlgorithmsACM Transactions on Information Systems 33 (3), 2015
Renseignements — A General SIMD-based Approach to Accelerating Compression Algorithms PDF (arXiv) — A General SIMD-based Approach to Accelerating Compression Algorithms
-
Adina Crainiceanu, Daniel LemireBloofi: Multidimensional Bloom FiltersInformation Systems 54, 2015
Renseignements — Bloofi: Multidimensional Bloom Filters PDF (arXiv) — Bloofi: Multidimensional Bloom Filters Code — Bloofi: Multidimensional Bloom Filters
-
Daniel Lemire, Leonid BoytsovDecoding billions of integers per second through vectorizationSoftware: Practice & Experience 45 (1), 2015
Renseignements — Decoding billions of integers per second through vectorization PDF (arXiv) — Decoding billions of integers per second through vectorization Diapositives — Decoding billions of integers per second through vectorization Code — Decoding billions of integers per second through vectorization
-
Antonio Badia, Daniel LemireFunctional dependencies with null markersComputer Journal 58 (5), 2015
Renseignements — Functional dependencies with null markers PDF (arXiv) — Functional dependencies with null markers
-
Xiaodan Zhu, Peter Turney, Daniel Lemire, Andre VellinoMeasuring academic influence: Not all citations are equalJournal of the Association for Information Science and Technology 66 (2), 2015
Renseignements — Measuring academic influence: Not all citations are equal PDF (arXiv) — Measuring academic influence: Not all citations are equal Jeu de données — Measuring academic influence: Not all citations are equal
-
Jeff Plaisance, Nathan Kurz, Daniel LemireVectorized VByte DecodingInternational Symposium on Web Algorithms 2015, 2015
Renseignements — Vectorized VByte Decoding PDF (arXiv) — Vectorized VByte Decoding Diapositives — Vectorized VByte Decoding Code — Vectorized VByte Decoding
-
Owen Kaser, Daniel LemireStrongly universal string hashing is fastComputer Journal 57 (11), 2014
Renseignements — Strongly universal string hashing is fast PDF (arXiv) — Strongly universal string hashing is fast Code — Strongly universal string hashing is fast
-
Hazel Webb, Owen Kaser, Daniel LemireDiamond DicingData & Knowledge Engineering 86, 2013
Renseignements — Diamond Dicing PDF (arXiv) — Diamond Dicing
-
Daniel Lemire, Owen Kaser, Eduardo GutarraReordering Rows for Better Compression: Beyond the Lexicographic OrderACM Transactions on Database Systems 37 (3), 2012
Renseignements — Reordering Rows for Better Compression: Beyond the Lexicographic Order PDF (arXiv) — Reordering Rows for Better Compression: Beyond the Lexicographic Order Diapositives — Reordering Rows for Better Compression: Beyond the Lexicographic Order Code — Reordering Rows for Better Compression: Beyond the Lexicographic Order Code 2 — Reordering Rows for Better Compression: Beyond the Lexicographic Order Code 3 — Reordering Rows for Better Compression: Beyond the Lexicographic Order
-
Daniel LemireThe universality of iterated hashing over variable-length stringsDiscrete Applied Mathematics 160 (4-5), 2012
Renseignements — The universality of iterated hashing over variable-length strings PDF (arXiv) — The universality of iterated hashing over variable-length strings
-
Zoltán Prekopcsák, Daniel LemireTime Series Classification by Class-Specific Mahalanobis DistancesAdvances in Data Analysis and Classification 6 (3), 2012
Renseignements — Time Series Classification by Class-Specific Mahalanobis Distances PDF (arXiv) — Time Series Classification by Class-Specific Mahalanobis Distances
-
Antonio Badia, Daniel LemireA Call to Arms: Revisiting Database DesignSIGMOD Record 40 (3), 2011
Renseignements — A Call to Arms: Revisiting Database Design PDF (arXiv) — A Call to Arms: Revisiting Database Design
-
Daniel Lemire, Andre VellinoExtracting, Transforming and Archiving Scientific DataIn VLDL 2011, Berlin, Germany, 2011
Renseignements — Extracting, Transforming and Archiving Scientific Data PDF (arXiv) — Extracting, Transforming and Archiving Scientific Data
-
Daniel Lemire, Owen KaserReordering Columns for Smaller IndexesInformation Sciences 181 (12), 2011
Renseignements — Reordering Columns for Smaller Indexes PDF (arXiv) — Reordering Columns for Smaller Indexes
-
Daniel Lemire, Owen KaserRecursive n-gram hashing is pairwise independent, at bestComputer Speech & Language 24 (4), pages 698-710, 2010
Renseignements — Recursive n-gram hashing is pairwise independent, at best PDF (arXiv) — Recursive n-gram hashing is pairwise independent, at best Code — Recursive n-gram hashing is pairwise independent, at best
-
Daniel Lemire, Owen Kaser, Kamel AouicheSorting improves word-aligned bitmap indexesData & Knowledge Engineering 69 (1), 2010
Renseignements — Sorting improves word-aligned bitmap indexes PDF (arXiv) — Sorting improves word-aligned bitmap indexes Code — Sorting improves word-aligned bitmap indexes
-
Daniel LemireFaster retrieval with a two-pass dynamic-time-warping lower boundPattern recognition 42 (9), 2009
Renseignements — Faster retrieval with a two-pass dynamic-time-warping lower bound PDF (arXiv) — Faster retrieval with a two-pass dynamic-time-warping lower bound Code — Faster retrieval with a two-pass dynamic-time-warping lower bound
-
Daniel Lemire, Martin Brooks, Yuhong YanAn Optimal Linear Time Algorithm for Quasi-Monotonic SegmentationInternational Journal of Computer Mathematics 86 (7), 2009
Renseignements — An Optimal Linear Time Algorithm for Quasi-Monotonic Segmentation PDF (arXiv) — An Optimal Linear Time Algorithm for Quasi-Monotonic Segmentation Code — An Optimal Linear Time Algorithm for Quasi-Monotonic Segmentation
-
Daniel Lemire, Owen KaserHierarchical Bin Buffering: Online Local Moments for Dynamic External Memory ArraysACM Transactions on Algorithms 4(1): 14 (2008)
Renseignements — Hierarchical Bin Buffering: Online Local Moments for Dynamic External Memory Arrays PDF (arXiv) — Hierarchical Bin Buffering: Online Local Moments for Dynamic External Memory Arrays Code — Hierarchical Bin Buffering: Online Local Moments for Dynamic External Memory Arrays
-
Owen Kaser, Daniel Lemire, Kamel AouicheHistogram-Aware Sorting for Enhanced Word-Aligned Compression in Bitmap IndexesDOLAP 2008
Renseignements — Histogram-Aware Sorting for Enhanced Word-Aligned Compression in Bitmap Indexes PDF (arXiv) — Histogram-Aware Sorting for Enhanced Word-Aligned Compression in Bitmap Indexes Code — Histogram-Aware Sorting for Enhanced Word-Aligned Compression in Bitmap Indexes
-
Hazel Webb, Owen Kaser, Daniel LemirePruning Attribute Values From Data Cubes with Diamond DicingIDEAS 2008
Renseignements — Pruning Attribute Values From Data Cubes with Diamond Dicing PDF (arXiv) — Pruning Attribute Values From Data Cubes with Diamond Dicing
-
Daniel LemireA Better Alternative to Piecewise Linear Time Series SegmentationSIAM Data Mining 2007
Renseignements — A Better Alternative to Piecewise Linear Time Series Segmentation PDF (arXiv) — A Better Alternative to Piecewise Linear Time Series Segmentation Code — A Better Alternative to Piecewise Linear Time Series Segmentation
-
Kamel Aouiche, Daniel LemireA Comparison of Five Probabilistic View-Size Estimation Techniques in OLAPDOLAP 2007, pp. 17-24, 2007
Renseignements — A Comparison of Five Probabilistic View-Size Estimation Techniques in OLAP PDF (arXiv) — A Comparison of Five Probabilistic View-Size Estimation Techniques in OLAP Diapositives — A Comparison of Five Probabilistic View-Size Estimation Techniques in OLAP Code — A Comparison of Five Probabilistic View-Size Estimation Techniques in OLAP
-
Dan Kucerovsky, Daniel LemireMonotonicity Analysis over Chains and CurvesIn Curves and Surfaces 2006, Saint-Malo, France, 2007
Renseignements — Monotonicity Analysis over Chains and Curves PDF (arXiv) — Monotonicity Analysis over Chains and Curves
-
Owen Kaser, Daniel LemireRemoving Manually-Generated Boilerplate from Electronic Texts: Experiments with Project Gutenberg e-BooksCASCON 2007
Renseignements — Removing Manually-Generated Boilerplate from Electronic Texts: Experiments with Project Gutenberg e-Books PDF (arXiv) — Removing Manually-Generated Boilerplate from Electronic Texts: Experiments with Project Gutenberg e-Books Code — Removing Manually-Generated Boilerplate from Electronic Texts: Experiments with Project Gutenberg e-Books
-
Owen Kaser, Daniel LemireTag-Cloud Drawing: Algorithms for Cloud VisualizationTagging and Metadata for Social Information Organization (WWW 2007)
Renseignements — Tag-Cloud Drawing: Algorithms for Cloud Visualization PDF (arXiv) — Tag-Cloud Drawing: Algorithms for Cloud Visualization Diapositives — Tag-Cloud Drawing: Algorithms for Cloud Visualization Code — Tag-Cloud Drawing: Algorithms for Cloud Visualization Jeu de données — Tag-Cloud Drawing: Algorithms for Cloud Visualization
-
Owen Kaser, Daniel LemireAttribute Value Reordering For Efficient Hybrid OLAPInformation Sciences 176 (16) 2006
Renseignements — Attribute Value Reordering For Efficient Hybrid OLAP PDF (arXiv) — Attribute Value Reordering For Efficient Hybrid OLAP
-
Daniel LemireStreaming Maximum-Minimum Filter Using No More than Three Comparisons per ElementNordic Journal of Computing 13 (4), pages 328-339, 2006
Renseignements — Streaming Maximum-Minimum Filter Using No More than Three Comparisons per Element PDF (arXiv) — Streaming Maximum-Minimum Filter Using No More than Three Comparisons per Element Code — Streaming Maximum-Minimum Filter Using No More than Three Comparisons per Element Code 2 — Streaming Maximum-Minimum Filter Using No More than Three Comparisons per Element
-
Daniel Lemire, Harold Boley, Sean McGrath, Marcel BallCollaborative filtering and inference rules for context‐aware learning object recommendationInteractive Technology and Smart Education 2 (3), 2005
Renseignements — Collaborative filtering and inference rules for context‐aware learning object recommendation PDF — Collaborative filtering and inference rules for context‐aware learning object recommendation
-
Daniel LemireScale and Translation Invariant Collaborative Filtering SystemsInformation Retrieval 8 (1), 2005
Renseignements — Scale and Translation Invariant Collaborative Filtering Systems PDF — Scale and Translation Invariant Collaborative Filtering Systems
-
Daniel Lemire, Anna MaclachlanSlope One Predictors for Online Rating-Based Collaborative FilteringIn SIAM Data Mining (SDM 2005), Newport Beach, California, April 21-23, 2005
Renseignements — Slope One Predictors for Online Rating-Based Collaborative Filtering PDF (arXiv) — Slope One Predictors for Online Rating-Based Collaborative Filtering
-
Daniel LemireA family of 4-point dyadic high resolution subdivision schemesIn Curves and Surfaces 2002, Saint-Malo, France, 2003
Renseignements — A family of 4-point dyadic high resolution subdivision schemes PDF — A family of 4-point dyadic high resolution subdivision schemes Code — A family of 4-point dyadic high resolution subdivision schemes
-
Serge Dubuc, Daniel Lemire, Jean-Louis MerrienFourier analysis of 2-point Hermite interpolatory subdivision schemesJournal of Fourier Analysis and Applications 7 (5), 2001
Renseignements — Fourier analysis of 2-point Hermite interpolatory subdivision schemes PDF — Fourier analysis of 2-point Hermite interpolatory subdivision schemes
-
Daniel Lemire, Chantal Pharand, Jean-Claude Rajaonah, Benoît Dubé, A.-Robert LeBlancWavelet time entropy, T wave morphology and myocardial ischemiaIEEE Transactions on Biomedical Engineering 47 (7), 2000
Renseignements — Wavelet time entropy, T wave morphology and myocardial ischemia
-
Gilles Deslauriers, Serge Dubuc, Daniel LemireUne famille d'ondelettes biorthogonales sur l'intervalle obtenue par un schéma d'interpolation itérativeAnnales des Sciences Mathématiques du Québec 23 (1), 1999
Renseignements — Une famille d'ondelettes biorthogonales sur l'intervalle obtenue par un schéma d'interpolation itérative PDF — Une famille d'ondelettes biorthogonales sur l'intervalle obtenue par un schéma d'interpolation itérative
Conférences
Je donne régulièrement des conférences. Ma conférence à QCon San Francisco 2019 a été identifiée comme “best voted” avec un taux de satisfaction de 98% ce qui est beaucoup plus élevé que la moyenne.
-
SIMD-Accelerated Data Processing
09/05/2026, SIMD-Accelerated Data Processing
-
Ada: Parsing Millions of URLs per Second
11/11/2023, NodeConf EU 2023
-
Binary Fuse Filters: Fast and Tiny Immutable Filters
16/06/2023, Invited talk at the Filter Workshop, Workshop held in conjunction with SPAA 2023 (June 16, 2023 - Orlando, USA)
-
Accurate and efficient software microbenchmarks
25/02/2023, Invited talk at the SIGPLAN BID 2023, Benchmarking in the Data Center: Expanding to the Cloud, Workshop held in conjunction with PPoPP 2023: Principles and Practice of Parallel Programming 2023 (February 25, 2023 - Montreal, Canada)
-
Accurate and efficient software microbenchmarks
25/02/2023, Invited talk at the SIGPLAN BID 2023, Benchmarking in the Data Center: Expanding to the Cloud, Workshop held in conjunction with PPoPP 2023: Principles and Practice of Parallel Programming 2023 (February 25, 2023 - Montreal, Canada)
-
Unicode at gigabytes per second
01/10/2021, Invited talk at SPIRE 2021, 28th International Symposium on String Processing and Information Retrieval (October 4-6th, 2021 - Lille, France)
-
Parsing numbers at a gigabyte per second
12/05/2021, MIT Fast Code Seminar
-
Floating-point Number Parsing with Perfect Accuracy at GB/s
07/10/2020, Go Systems
-
Data Engineering at the Speed of Your Disk
16/06/2020, Performance Summit III (Facebook)
-
Parsing JSON Really Quickly: Lessons Learned
07/10/2019, QCon San Francisco 2019
Projets
fastfloat
Routines rapides de lectures de nombres à virgule
SIMDJSON
Traiter des gigaoctets de documents JSON par seconde
SIMDUTF
Routines Unicode : des milliards de caractères par seconde
Les bitmaps Roaring
Bitmap compressés et véloces, largement déployés. (photo: Edge Earth)
Laboratoire
Nous avons la chance d’avoir un laboratoire entièrement équipé avec un technicien dédié. Nous disposons d’une ferme de serveurs qui a été utilisée dans le monde entier pour des expériences sur la performance des logiciels (par exemple, par des chercheurs comme Agner Fog). Nous disposons également de plusieurs stations de travail puissantes et de magnifiques tableaux blancs !
Enseignement
Premier cycle :
- INF 1220 - Introduction à la programmation
- INF 2007 - Programmation avancée
- INF 2020 - Programmation d’applications avec Python : des jeux au Web
- INF 4450 - Programmation orientée-données
- INF 6460 - Recherche et filtrage d’informations
- INF 9004 - Informatique des entrepôts de données
Cycle supérieur :
- INF 6104 - Recherche d’informations et Web
- INF 6107 - Web social
- INF 6408 - Informatique de l’analyse multidimensionnelle
Programmes :
Étudiants
Quelques diplômés récents, par diplôme et du plus récent au plus ancien :
Doctorat
Maîtrise avec mémoire
Maîtrise professionnelle (sans mémoire)
Une sélection d’étudiants.
| Année | Diplômé |
|---|---|
| 2025 | Antoine Tohme |
| 2024 | Ernso Decelien |
| 2023 | Zakia Chaibeddra bourse Alexander-Graham-Bell |
| 2022 | Nicolas Boulet-Lavoie |
| 2020 | Achille Alain Kamdem Djoko développeur chez Bell Canada |
| 2020 | Hakim Berraki concepteur de solutions à la Société de transport de Montréal |
| 2020 | Abdelkader Kaddour Brahim |
| 2018 | Massamba Fall analyste principal chez Promutuel · GitHub |
| 2017 | Maxime Boisvert directeur de l’ingénierie chez Shopify |
| 2017 | Shany Carle professeur d’informatique au cégep de Victoriaville |

Quelques ancients étudiants:
- Luis Garcia Vargas
- Shira Smith travaille comme ingénieure chez Discord en Californie.
- Dara Aghamirkarimi enseigne l’informatique au collège LaSalle.
- Geneviève Lefebvre (décédée du cancer pendant ses études de maîtrise)
Étudiants au doctorat en cours de supervision:
- Guy Jobin (dirigé avec Dragos Vieru)
- Khargou Jalal
- Sofiane Faïdi
- Faten Slama
- William Ouedraogo
- Jean-Vincent Bogui
- Aubrey Trask
- Ali Lienaux
- Papa Waly Diouf (dirigé avec Richard Hotte)
Étudiants à la maîtrise en cours de supervision:
- Isaac Hurtubise
- Juan Hernandez
- Nicolas Irep
- Emna Ben Hamouda
- Aaron Rivera
- Khalid Bouraki
- Yan Levasseur
Stagiaires post-doctoraux récents:
- Jaël Champagne-Gareau (2025–2027) dblp
Assistants de recherche récents (premier cycle):
- Nick Nuon, été 2023 et 2024, récipiendaire d’une bourse de recherche de premier cycle du CRSNG.
- Nicolas Boyer, été 2021 et 2022, récipiendaire d’une bourse de recherche de premier cycle du CRSNG. Nicolas est développeur logiciel chez Avid.
- David Favreau, automne 2021.
- Yoann Le Rouzic, été 2020. GitHub
- Io Andes Daza-Dillon, été 2019, récipiendaire d’une bourse de recherche de premier cycle du CRSNG. Io est consultant chez Savoir-faire Linux. GitHub
- Jérémie Piotte, automnes 2018 et 2019, récipiendaire d’une bourse de recherche de premier cycle du CRSNG. Jérémie est Senior Manager (Machine Learning Engineering) chez Unity Technologies GitHub
- Niko Girardelli, hiver 2018. GitHub
Invités de recherche récents :
Mentorat
- Homma Kazutaka, Google Summer of Code, été 2018 (co-mentor avec Harlan Haskins chez Apple)
Nouvelles
Bérenger Bramas a passé son HDR !
Bérenger Bramas a passé son HDR.
Guy Carlos Tamkodjou Tchio est un nouveau docteur !
Guy Carlos Tamkodjou Tchio est un nouveau docteur !
CONTINUER DE LIRE — Guy Carlos Tamkodjou Tchio est un nouveau docteur !
Lockman Saleh reçoit son doctorat!
Lockman Saleh a soutenu avec succès sa thèse de doctorat.
Fatma Miladi est une nouvelle docteure !
Fatma Miladi a soutenu avec succès sa thèse de doctorat.
CONTINUER DE LIRE — Fatma Miladi est une nouvelle docteure !
Services
Je suis éditeur de la revue Software: Practice and Experience (Wiley) depuis 2021. Cette revue a été fondée en 1971 et elle a publié plusieurs articles fondamentaux en informatique.
Avant les événements de 2020, j’organisais à Montréal des séries d’ateliers ouverts au public: le technolab et le tribalab. En 2019, j’ai été le président d’EDA 2019 (Business Intelligence & Big Data) tenue en octobre 2019 à Montpellier, France. En juin 2018, j’ai participé au séminaire Dagstuhl 18251 intitulé “Database Architectures for Modern Hardware”. En 2018, j’ai été reconnu par la revue Software: Practice and Experience comme “distinguished referee”. J’ai été éditeur associé de la section informatique au sein de la revue Heliyon (Elsevier) de 2015 à 2020.
J’ai récemment fait partie des comités scientifiques suivants :
- WSDM 2027: The 20th International Conference on Web Search and Data Mining (February 15-17, 2027, in Hong Kong, China) – Senior Member
- CIKM 2026: 34th ACM International Conference on Information and Knowledge (November 7-11, 2026 in Rome Italy) – Senior Member
- ECMLPKDD 2026: European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (September 7-11, 2026 in Naples Italy)
- WSDM 2026: The Nineteenth International Conference on Web Search and Data Mining (March 10-14, 2026, in Boise, Idaho, USA) – Senior Member
- CIKM 2025: 33rd ACM International Conference on Information and Knowledge (October 21-25, 2025 in Boise, Idaho) – Senior Member
- BIGDACI 2025: 10th International Conference on Big Data Analytics, Data Mining and Computational Intelligence (23-25 July 2025 in Lisbon, Portugal)
- SIGIR 2025: The 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (13-18 July 2025 in Padua, Italy)
- ECMLPKDD 2025: European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (15-19 September 2025, in Oporto, Portugal)
- WSDM 2025: The Eighteenth International Conference on Web Search and Data Mining (March 10-14, 2025, in Hannover, Germany) – Senior Member
- ECMLPKDD 2024: European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (September 9-13, 2023, in Vilnius, Italy)
- SIGIR 2024: The 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (Washington D.C., USA, 14-18 July, 2024).
- ECMLPKDD 2023: European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (September 18-22, 2023, in Turin, Italy)
- SIGKDD 2023: 29th SIGKDD Conference on Knowledge Discovery and Data Mining (Long Beach, California, August 6 2023)
- SIGIR 2023: The 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taiwan, July 23rd to 27th, 2023).
- EDA 2022: 18e journées EDA Business Intelligence and Big Data (Clermont-Ferrand,France, 27-28 octobre 2022)
- SIGIR 2022: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain, July 11-15, 2022)
- DOLAP 2022: 24th International Workshop on Design, Optimization, Languages and Analytical Processing of Big Data
- WSDM 2022 15th ACM International WSDM Conference (Phoenix, AZ, USA, Feb. 2nd to March 4th, 2022)
- ASD 2021: 13th edition of the Conference on Advances in the Science of Data (Blida, Algeria, 24-25 October 2021)
- CIKM 2021: 30th ACM International Conference on Information and Knowledge (Gold Coast, Queensland, Australia, 1-5 November 2021)
- ECML/PKDD 21: European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (Bilbao, Spain, 13-17 September)
- EDA 2021: 17e journées EDA Business Intelligence and Big Data (1-2 July 2021)
- SIGKDD 2021: 27th International Conference on Knowledge Discovery and Data Mining (Singapore, Aug 14-18, 2021)
- SIGIR 2021: 44th International ACM SIGIR Conference on Research and Development in Information Retrieval
- DOLAP 2021: 23nd International Workshop On Design, Optimization, Languages and Analytical Processing of Big Data
- WSDM 2021:14th ACM International WSDM Conference (Jerusalem, Israel, March 8-12, 2021)
- EDML20: Second Workshop on Evaluation and Experimental Design
- RecSys 2020: 14th ACM Recommender Systems Conference (Rio de Janeiro, Brazil)
- BBIGAP'2020: Second International Workshop for Business Intelligence & Big Data Applications
- ECML-PKDD 2020: European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (Area Chair)
- CIKM 2020: 29th ACM International Conference on Information and Knowledge
- DaWak 2020: 22nd International Conference on Big Data Analytics and Knowledge Discovery
- SIGIR 2020: 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval
- DOLAP 2020: 22nd International Workshop On Design, Optimization, Languages and Analytical Processing of Big Data
- ADBIS 2020: 24th European Conference on Advances in Databases and Information System (August 25-28: Lyon, France)
- CIKM 2019: 28th ACM International Conference on Information and Knowledge Management (Nov 3-7, 2019: Beijing, China)
- RecSys 2019: 13th ACM Recommender Systems Conference (September 2019; Copenhagen, Denmark)
- BBigAP 2019: 1st International Workshop on BI & Big Data applications, In conjunction with the 23rd European Conference on Advances in Databases and Information Systems (ADBIS 2019) (September 8, 2019; Bled, Slovenia)
- DOLAP 2019 - 21st International Workshop On Design, Optimization, Languages and Analytical Processing of Big Data (March 26, 2019; Lisbon, Portugal)
- CIKM 2018 - Twenty-Seventh ACM International Conference on Information and Knowledge Management (October 22-26, 2018; Turing, Italy)
- ASD 2018 - 12th edition of the Conference on Advances of Decisional Systems : Big data & Applications (May 2018; Marrakech, Morocco)
- RecSys 2018 - 12th ACM Recommender Systems Conference (October 2018; Vancouver, Canada)
- WABiD* 2018 - 1st International Workshop on Advances on Big Data Management, Analytics and Security (September 2018; Budapest, Hungary)
- WWW 2018 - Twenty-seventh International WWW Conference (April 23-27 2018; Lyon, France)
- DOLAP 2018 - Nineteenth International Workshop On Design, Optimization, Languages and Analytical Processing of Big Data (March 26–29, 2018; Vienna, Austria)
- CIKM 2017 - Twenty-Sixth ACM International Conference on Information and Knowledge Management (November 6-10, 2017; Singapore)
- SPIRE 2017 - 24th International Symposium on String Processing and Information Retrieval (September 26-29, 2017; Palermo, Italy)
J’ai été un examinateur externe sur les thèses de doctorat suivantes :
- Azim Afroozeh à la Vrije Universiteit Amsterdam (2026) – dirigé par Peter Boncz.
- Lockman Saleh à l’UQAM (2025) - dirigé par Hafedh Mili et Mounir Boukadoum.
- Jaël Champagne Gareau à l’UQAM (2024) - dirigé par Éric Beaudry.
- Nathan Maurice à la Sorbonne, France (2024) - dirigé par Lionel Lacassagne.
- Nigel Medforth à l’Université Simon Fraser (2022) - dirigé par Robert Cameron.
- Luca Versari à l’Université de Pise (2021) - dirigé par Roberto Grossi.
- Kareem El Gebaly à l’Université Waterloo (2018) - dirigé par Jimmy Lin, Lukasz Golab et Ashraf Aboulnaga.
- Mohammed Shaaban à l’Université Pierre et Marie Curie (2017) - dirigé par Patrick Garda.
- Mehdi Boukhechba à l’UQAC (2016) - dirigé par Abdenour Bouzouane et Charles Gouin-Vallerand.
- Hicham Assoudi à l’UQAM (2016) - dirigé par Hakim Lounis.
- Khaled Dehdouh à Lyon 2 (2015) - dirigé par Omar Boussaid.
- Martin Leginus à l’Université Aalborg (2015) - dirigé par Peter Dolog.
- Ahmad Taleb à l’Université Concordia (2011) - dirigé par Todd Eavis.
J’ai évalué les mémoires de maîtrise suivants:
- Benjamin Lapointe-Pinel de l’UQAR, Canada (2024) - dirigé par Steven Pigeon.
En 2020, j’étais l’un de deux évaluateurs externes du programme de maîtrise en informatique à l’UQAC.
J’ai servi comme membre de comité d’évaluateur au sein d’organismes subventionnaires :
- FRQNT: comité d’évaluation 03F (informatique théorique) depuis 2007.
- FRQNT: comité d’évaluation 309 (subvention d’équipe en informatique) depuis 2006.
- CRSNG: comité d’évaluation du programme de subventions d’outils et d’instruments de recherche dans les sciences informatiques (2012-2015)
- CRSNG: comité d’évaluation des subventions à la découverte en Sciences informatiques, comité 1507 (2018-2021), co-président du comité en 2019-2020 et 2020-2021.
- CRSNG: comité d’évaluation de Horizons de la découverte (2022)
En 2022, j’ai fait partie du sous-comité universitaire sur le génie et les technologies de l’information, au sein du comité sur l’implantation des mesures de l’opération main-d’oeuvre du gouvernement du Québec.
Livres
Java pas à pas
Programmation avec Python: des jeux au Web
La science des données: Théorie et applications avec R et Python
Maîtriser la programmation: Des tests à la performance en Go
Mastering Programming: From Testing to Performance in Go
Faster Than You Think: Essays on Thinking Better and Building Faster
Média
Articles et entrevues
- Algorithmes et réseaux sociaux, CHOI FM 93 (Radio X), 9 février 2026.
- Des milliards pour les éoliennes au Québec: pourquoi et pour qui?, Libre Média, 15 septembre 2025.
- Pour un Québec prospère et une approche équilibrée en environnement, Libre Média, 3 juin 2025.
- Trump, nouvel alibi des Libéraux pour gouverner par la peur, Libre Média, 2 avril 2025.
- On SIMD, cache and CPU internals with the expert Daniel Lemire!, Wookash Podcast, 21 février 2025.
- Avenir de la science (entrevue), Libre Média, 18 février 2025.
- Les géants du numérique raffolent des algorithmes de ce prof québécois, Journal de Montréal, 19 octobre 2024.
- Le professeur Daniel Lemire de l’Université TÉLUQ parmi les chercheurs les plus cités au monde, CNW/Université TÉLUQ, 15 octobre 2024.
- Artificial Intelligence Is the Crisis We Need, Communications of the ACM (blog), 6 juin 2024.
- The Service of Dissent, Brownstone Institute, 3 mai 2024.
- Les universités de Québec ont « l’oreille tendue » vers la recherche sur l’IA, Radio-Canada, 24 avril 2024.
- L’intelligence artificielle à l’Université TÉLUQ, Téléjournal ICI Québec, 23 avril 2024.
- Un robot conversationnel dans certains cours de la TÉLUQ comme outil d’aide pédagogique, Journal de Québec, 28 février 2024.
- Making Parsing I/O Bound with Daniel Lemire, Software Unscripted, 17 août 2023.
- ¿Supone ChatGPT el fin de los programadores? “Los hará más eficientes en lugar de reemplazarlos”, El País (9 juin 2023)
- ChatGPT is Not a Technological Singularity, Communications of the ACM (blog), 5 juin 2023. (Un des billets les plus lus de 2023 chez ACM.)
- Identité numérique: Une solution, mille interrogations, La Presse, 21 novembre 2022.
- Liberté académique: l’Université Laval doit s’excuser et réparer ses torts, lettre d’opinion dans Le Soleil, 15 août 2022.
- L’identité numérique: quels enjeux et quels risques?, lettre d’opinion dans Le Journal de Montréal, 10 juin 2022.
- Ajouter sa brique à l’édifice de Microsoft, Radio-Canada (Moteur de recherche), 6 mai 2022.
- Microsoft utilise le mémoire universitaire d’un étudiant pour améliorer ses fonctionnalités, Figaro, 11 mars 2022.
- Microsoft va utiliser les travaux d’un étudiant à la maîtrise, Les Affaires, 9 mars 2022.
- Qu’attendez-vous pour adopter le «no-code»?, Les Affaires, 31 janvier 2022.
- Protéger et promouvoir les libertés académiques pour tous, lettre d’opinion dans Le Soleil, 25 septembre 2021.
- Commission parlementaire québécoise sur la liberté académique (CTV News), court reportage, 24 août 2021.
- TwitterSpaces with Daniel Lemire (interview), Denis Bakhalov, 3 mai 2021.
- Mémoire adressé à la commission scientifique et technique indépendante sur la reconnaissance de la liberté académique dans le milieu universitaire, juin 2021.
- Frontiers of Performance with Daniel Lemire (entrevue), Corecursive, 1er décembre 2020.
- The research paper should NOT be the final product (entrevue), A.I. Socratic Circles, 3 mars 2020.
- Optimizations in C++ Compilers by Godbolt (cite mon blogue). Communications of the ACM, February 2020.
- Des serveurs informatiques plus rapides et moins énergivores (entrevue), Les années lumières, Radio-Canada, 13 décembre 2019.
- Des langues découpées en bits pour comparer leur efficacité (entrevue), Les années lumières, Radio-Canada, 15 septembre 2019.
- L’intelligence artificielle pour rendre les logiciels plus rapides (article de magazine sur les travaux de Daniel Lemire), Québec Science, 28 mars 2019.
- Défendre la liberté académique, lettre d’opinion dans Le Soleil, 29 janvier 2019.
- Open Source Powers Supercomputing by Shein (cite mon point de vue). Communications of the ACM, February 2018.
- Uses This: Daniel Lemire (entrevue), Uses This, 5 novembre 2013.
Conseil
J’ai travaillé comme consultant depuis 1998. En tant que consultant, j’ai construit des logiciels personnalisés, j’ai résolu des problèmes de performance profonds, j’ai offert des sessions de formation spécialisées, j’ai conçu des algorithmes novateurs. J’adore travailler avec l’industrie sur des problèmes importants. Je propose les services commerciaux suivants.
- Conférences. Je suis ravi de parler à votre équipe des avancées récentes en logiciel ou d’autres sujets pertinents. Mes conférences sont bien accueillies. Mes tarifs varient de 5 000 $ à 15 000 $ par engagement, en fonction de la durée, du format (virtuel ou en personne) et des frais de déplacement, qui sont généralement couverts séparément.
- Formation. Je fournis une formation avancée exclusive et de haute qualité pour votre équipe. Pour les sessions de formation d’entreprise en ingénierie logicielle, mes tarifs se situent entre 3 000 $ et 5 000 $ par jour, en fonction de la personnalisation, de la durée (demi-journée ou journée complète) et du nombre de participants.
- Consultation. Si vous avez des problèmes spécifiques dans votre entreprise, je serai ravi de venir vous aider. Mes honoraires de consultation sont de 400 $ par heure, avec des frais supplémentaires pour les déplacements. Lorsque c’est possible, je préfère proposer des services à forfait fixe (par exemple, 5 000 $).
- Logiciels open-source sponsorisés. Certains de mes travaux open-source ont été sponsorisés par des entreprises privées et développés selon leurs besoins. Les sponsorisations pour les projets open-source peuvent être structurées en niveaux mensuels ou en contributions uniques pour des fonctionnalités spécifiques, allant de 1 000 $ à 10 000 $ ou plus, en fonction de l’ampleur. J’encourage les entreprises à me sponsoriser sur GitHub, surtout si elles bénéficient de mon travail. Les niveaux de sponsorisation supérieurs sur GitHub sont destinés aux entreprises et incluent des avantages spécifiques.
Me joindre
- [email protected]
- Université du Québec (TÉLUQ), 5800, rue Saint-Denis, Bureau 1105, Montréal (Québec) H2S 3L5 Canada
- sur rendez-vous







