11 Aug 2026

Deep Learning Indaba 2026 Reflection: Days 5 & 6 (Workshop Days)

Image: Dr Idris Abdulmumin at the Deep Learning Indaba 2026 in Lagos, Nigeria.

Image: Dr Idris Abdulmumin at the Deep Learning Indaba 2026 in Lagos, Nigeria.

Dr Idris Abdulmumin, a Research Associate at DSFSI, co-organised two workshops at the Deep Learning Indaba 2026 in Lagos, Nigeria. Here are his notes from the workshop days.


The workshop days were the most useful part of the Indaba for me this year.

The Masakhane Workshop

The Cultivating Sovereign African NLP workshop brought together a panel that discussed the vast amounts of data African media organisations hold. One organisation in Kano alone has more than 16 TB spanning five years, and there are over 800 licenced media organisations producing content in Hausa, Igbo, Yoruba and Pidgin, as well as several other languages like Fulfulde, Kanuri, Tiv and Ijaw. A recurring point was the need for a collaboration mechanism that enables ethical access to and use of these archives for research and product development.

In my presentation, I walked the audience through using OCR and ASR to convert scanned documents and audio files into text, and the processes in between. The audience followed up with a suggestion to organise a shared task on OCR for African and other low-resourced languages, and one person mentioned ongoing work in East Africa on extracting medical records that we might collaborate on. There was also a licence panel on how and on what a person can claim a licence, and on the importance of consent, documentation and IP.

At the Masakhane dinner, Kathleen mentioned her work on tokenization at the syllable level. I followed up with her at the workshop and we discussed testing its efficacy on other Swahili dialects and wider African languages, using SLMs to validate.

The Tool and Data Playbook Workshop

The Data, Culture, and Community exhibition of the AfricaNLP Playbook and Annotation Tool was a 1.5-hour interactive session held at Pan-Atlantic University, funded by the Masakhane African Languages Hub. I co-organised it with Seid Muhie Yimam (University of Hamburg / Bahir Dar University), Shamsuddeen Hassan Muhammad (Bayero University Kano), Abinew Ali Ayele (Bahir Dar University), Tadesse Destaw Belay (Instituto Politécnico Nacional), and Ibrahim Said Ahmad (University of Wisconsin–Stevens Point).

At the session, we presented the playbooks we are creating, framed around the foundational limitations of the African data landscape: scarcity of data, the skew towards publishing models and reusing existing datasets rather than building new ones, and the lack of end-to-end documentation to guide a data creator from inception to publication of a dataset, including the legal and logistical processes that are supposed to be involved.

We closed with a live demo of the tool. The audience gave several recommendations, including data visitation (local data, online annotations), chunked uploads for large data, and sign language annotation support. We have since acted on the last one by building a new VideoRecord tag that enables gesture capture.

Since the workshop we have seen a lot of interaction with the annotation tool, and several people have written LinkedIn DMs about ways to use and collaborate on it.

A Footnote

It now seems that everyone knows me and I also know everyone. Maybe because this is my 4th Indaba?


Follow our journey at Deep Learning Indaba 2026 and beyond: