Jackson Cionek
3 Views

Big Data Is Not the World

Big Data Is Not the World

How reality becomes data - and what disappears before an AI even begins to learn

Imagine two people standing in front of the same river.

One arrived that morning. The other was born there. They recognize when the water smells different, remember where the soil gives way after heavy rain, know which fish used to appear at certain times of the year, and associate that place with work, stories, fear, belonging, and survival.

A camera may capture almost the same image for both of them.

But the image contains neither person's experience.

This is a useful starting point for thinking about Artificial Intelligence. Before a model can learn anything about the world, some part of that world must first be perceived, selected, measured, transduced, classified, and transformed into something computationally accessible.

That is why:

Big Data is not the world. It is an enormous collection of fragments that managed to become data.

Before data, there is a cut

A camera transforms light into signals. A microphone converts variations in air pressure. A satellite measures specific bands of the electromagnetic spectrum. A medical record translates part of a person's trajectory into fields, codes, and text. A platform transforms clicks, searches, viewing time, and shares into behavioral traces.

We can imagine the process as:

reality → cut → transduction → encoding → data → digital database → model → inference

The problem is not that a cut exists.

Without selecting aspects of reality, we would have no maps, photographs, statistics, medical records, EEG, fNIRS, or science.

The problem begins when we forget that a cut has taken place.

Mexican-Ecuadorian researcher Paola Ricaurte draws attention to this issue by approaching AI systems as sociotechnical assemblages shaped by institutions, interests, power relations, and particular ways of producing knowledge. The algorithmic problem, therefore, does not begin only when a dataset contains “bias.” It can begin within the structures that determine what will be measured, preserved, classified, and converted into information.

This forces us to ask a question even before asking anything about AI:

Who had the conditions to turn their reality into data?

What if we are looking for the key only where there is light?

Mexican anthropologist Renée de la Torre, while discussing Néstor García Canclini's Ciudadanos reemplazados por algoritmos (Citizens Replaced by Algorithms), retrieves a small story that is particularly useful here.

A man is searching for his keys under a streetlamp. Someone asks whether he lost them there. No, he replies. He dropped them near a tree several meters away.

Then why search under the streetlamp?

Because that is where the light is.

It is a powerful metaphor for Big Data.

We may possess millions of records on income, purchases, credit, mobility, production, searches, consumption, and social-media behavior because we have built sensors, platforms, and institutions capable of continuously illuminating those dimensions.

At the same time, a territory may contain far fewer records on locally perceived water quality, changes in species, traditional knowledge, community ties, belonging, or gradual transformations of a Biome as perceived by the people who actually live there.

Those phenomena do not have to be incorrectly recorded.

They may simply not be recorded at all.

Perhaps we are searching for reality precisely where our databases are best able to illuminate it.

This is what we call epistemic invisibility: what fails to enter the representational universe may also fail to participate in subsequent inferences and decisions.

Big Data can mean a great deal of information about a small part of the lived world.

García Canclini himself has shown how opinions and behaviors transformed into data can subsequently be reorganized by corporations and algorithms, reshaping forms of citizenship and participation.

The sign is not the whole object

Semiotics allows us to avoid two extremes.

We do not need a representation to contain all of reality. A map works precisely because it reduces.

But we cannot forget what the map leaves out.

Anderson Vinícius Romanini has been bringing Peircean semiosis into dialogue with Active Inference. In a 2025 publication, he discusses intelligent processes as inferential processes capable of navigating environments, dealing with uncertainty, and stabilizing habits.

Together with Marcelo Hamdan Alvim, Romanini also examines how algorithmic mediation can participate in users' own inferential processes, shaping beliefs and directing patterns of interaction.

This means that data do not simply remain stored.

They return.

They are organized by algorithms, become recommendations, responses, rankings, or decisions, and re-enter the perceptual environment of the Body-Territory.

Maria Eunice Quilici Gonzalez and colleagues add another important layer. In discussing autonomy in the age of Big Data, they argue that people may remain capable of relatively rational decisions while simultaneously having their opinions influenced by insufficient or distorted information, habits, and previously acquired emotional dispositions.

Therefore, it is not enough to ask:

Is the inference correct in relation to the available data?

We must also ask:

Do the available data sufficiently represent the world upon which that inference intends to act?

The screen transduces. The Body-Territory lives.

At BrainLatam, we use 12D as a perceptual cartography of the Body-Territory.

We are not claiming that neuroscience recognizes exactly twelve independent senses, nor that the physical universe has twelve dimensions.

The purpose is to remind us that being in the world involves much more than vision and hearing: proprioception, interoception, balance, nociception, thermosensation, internal chemical signals, movement, and multiple bodily and environmental relations. In the BrainLatam formulation, belonging remains a hypothesis to be investigated, not an established sensory modality.

A forest displayed in 8K can be extraordinarily rich visually.

But the screen does not contain a smaller piece of that forest.

It creates another material configuration — pixels, light, and sound — capable of representing certain relationships present in the original event.

In another BrainLatam article, we used the metaphor “The Logos Crystallized into AI” to think about how part of humanity's linguistic and logical-formal production has been incorporated into computational systems.

Now we can move one step further back:

Before the Logos crystallized into AI, the world had to crystallize into data.

And not every world had the same opportunity to do so.

Neuroscience also cuts reality

The same caution must apply to our own instruments.

EEG does not measure “the person,” nor does it directly read thoughts. It records differences in electrical potential at the scalp, from which temporal, spectral, and spatial properties related to neural activity can be extracted.

fNIRS does not measure “consciousness.” It uses near-infrared light to monitor hemodynamic changes, including variations associated with oxygenated and deoxygenated hemoglobin concentrations in cortical regions accessible to the technique.

They are extraordinary windows.

But they remain windows.

Recognizing this does not diminish EEG or NIRS. On the contrary, it prevents the physiological record from being confused with the entire phenomenon.

In 2025, Faisal Mushtaq and Agustín Ibáñez argued for the potential of EEG to support a genuinely global neuroscience. Its portability, relatively lower cost, and scalability can help include populations historically underrepresented in neuroscience. At the same time, the authors emphasize the need for cross-site harmonization, hardware suited to human diversity, community participation, and culturally sensitive methodologies.

Here, Decolonial Neuroscience ceases to be merely a conceptual critique.

It enters the method itself.

Taking a portable EEG system into a community does not automatically make research decolonial. If the scientific question was formulated far from that territory, the categories arrived ready-made, the data leave the community and never return, and only outside researchers retain the authority to interpret them, we may simply have made an old asymmetry portable.

Territory can participate in the signal without being “written” in it

A study published in Nature Medicine in 2024 helps make this issue quantitatively visible.

The research analyzed 5,306 participants across 15 countries, seven of them in Latin America and the Caribbean, using EEG and fMRI. The study examined differences between estimated brain age and chronological age and found associations with geographic diversity and socioeconomic, sociodemographic, and environmental disparities.

The international team included substantial Latin American participation, with researchers such as Sebastián Moguilner, Sandra Báez, Pedro Valdés-Sosa, Francisco Lopera, and Agustín Ibáñez.

This does not mean that we can look at an EEG and “read” someone's territory.

It means something more careful:

The brain that arrives at the recording system already has a history.

Education, health, nutrition, pollution, inequality, culture, and environment have participated in that trajectory before the first electrode is placed.

Decolonial Neuroscience should therefore not search for a supposed “Latin American brain.”

It should ask:

How much of what we classify as individual difference also carries territorial histories that our experimental protocol did not measure?

EEG and NIRS can help bring neuroscience back into the world

The technologies themselves are opening new possibilities.

In 2023, Ernesto Vidal-Rosas and colleagues highlighted how wearable and high-density fNIRS systems can bring functional neuroimaging closer to environments and behaviors resembling everyday life.

In Brazil, Priscila Benitez and colleagues used fNIRS in a naturalistic clinical setting during educational activities involving children and young people with autism and/or intellectual disability, demonstrating the feasibility of moving recordings toward tasks that are less detached from lived situations.

This may be an especially important opportunity for Latin America.

We can use EEG and fNIRS to produce one quantifiable layer of the Body-Territory, carefully relating it to behavior, first-person experience, culture, and environment — without pretending that physiological signals are equivalent to the experience as a whole.

Science becomes stronger when it knows where its instrument ends.

NeuroDesafio LATAM research question

As a sponsor of NeuroDesafio LATAM / NeuroChallenge LATAM, BrainLatam proposes turning this discussion into an experimental question:

How do EEG and fNIRS responses to the same digital stimulus change when that stimulus is presented to Body-Territories living in different cultural, social, and environmental contexts across Latin America?

A multicenter study could present identical digital stimuli — including AI-generated stimuli — in different Latin American regions, combining EEG and/or fNIRS with behavior, first-person reports, and territorial variables.

The question would not be:

“Which population has a different brain?”

It would be:

How much of the response we call individual depends on the conditions in which that Body-Territory learned to perceive, signify, and act?

Perhaps this is where Big Data, Semiotics, Artificial Intelligence, and Decolonial Neuroscience can meet most productively.

We do not need to put the entire world inside a database.

That would be impossible.

We need to build science — and intelligences — capable of recognizing that:

what they can represent is never the whole world.


Commented References — a second reading of the article

  1. Ricaurte Quijano, P. (2025). Algorithmic Assemblages of Power: AI Harm and the Question of Responsibility. Teknokultura.
    Ricaurte moves the discussion beyond the narrow notion of “algorithmic bias” toward the broader sociotechnical structures that produce data, categories, infrastructures, and decisions. Her work supports this article's first question: who had the power and infrastructure to turn their reality into data, and which relations of power determined what would become visible to AI?

  2. De la Torre Castellanos, A. R. (2021). “Ciudadanos reemplazados por algoritmos, Néstor García Canclini”. Alteridades, 31(62).
    The Mexican anthropologist recalls the metaphor of searching for lost keys under a streetlamp because that is where there is light. It condenses one of this article's central ideas: science and AI may achieve extremely high resolution where measurement infrastructure already exists while crucial aspects of reality remain outside the illuminated field. It is an accessible metaphor for what we call epistemic invisibility.

  3. García Canclini, N. (2019). Ciudadanos reemplazados por algoritmos. Universidad de Guadalajara / CALAS / FLACSO Ecuador.
    The Argentine-Mexican anthropologist examines how opinions and behaviors captured as data become embedded in the power of platforms and corporations. His work helps distinguish making citizens digitally legible from actually expanding citizenship: greater data visibility may increase the capacity of systems to observe individuals without proportionally increasing individuals' ability to observe or contest those systems.

  4. Gonzalez, M. E. Q.; Broens, M. C.; Quilici-Gonzalez, J. A.; Kobayashi, G. (2023). “Hábitos e racionalidade: um estudo filosófico-interdisciplinar sobre autonomia na era dos Big Data”. Trans/Form/Ação, 46, 367–386.
    The authors show that autonomy does not simply disappear under digital systems. People may act rationally while their opinions are still modulated by insufficient or distorted information, prior habits, and emotional dispositions. For this article, the publication is essential because it shows that the quality of a decision depends not only on false information but also on what is absent from the informational field.

  5. Romanini, A. V. (2025). “Semiose, inteligência e inferência ativa”. deSignis, 43, 61–72.
    Romanini brings Peircean semiotics into dialogue with Active Inference, creating a bridge among Semiotics, Cognitive Science, and AI. It supports our argument that signs are not merely stored representations; they participate in inferential processes through which systems navigate environments and stabilize ways of acting.

  6. Alvim, M. H.; Romanini, A. V. (2025). “Mediações algorítmicas e cognição: conexões entre a semiótica e a inferência ativa”. Esferas, 32.
    This work directly links algorithms, cognition, and social networks, discussing how algorithmic mediation can shape beliefs and response patterns. It completes the recursive circuit described here: the world becomes data, data feed algorithms, and algorithms return as new signs capable of modifying the behavior that will generate subsequent data.

  7. Moguilner, S.; Báez, S.; Hernandez, H. et al. (2024). “Brain clocks capture diversity and disparities in aging and dementia across geographically diverse populations”. Nature Medicine, 30, 3646–3657.
    Using EEG and fMRI data from 5,306 participants across 15 countries, the study links geographic diversity and socioeconomic, sociodemographic, and environmental disparities to brain-age measures. It provides important quantitative support for the Body-Territory perspective: the brain measured by an instrument arrives at the laboratory after a long social and environmental trajectory.

  8. Mushtaq, F.; Ibáñez, A. (2025). “Electroencephalography (EEG) and the Quest for an Inclusive and Global Neuroscience”. European Journal of Neuroscience, 61(6), e70078.
    The authors present EEG as an especially promising technology for global neuroscience because of its portability, accessibility, and scalability, while stressing hardware diversity, harmonization, and culturally co-produced methods. Their argument helps define Decolonial Neuroscience as a transformation in how neuroscientific data are produced, rather than merely a change in the discourse surrounding those data.

  9. Vidal-Rosas, E. E.; von Lühmann, A.; Pinti, P.; Cooper, R. J. (2023). “Wearable, high-density fNIRS and diffuse optical tomography technologies: a perspective”. Neurophotonics, 10(2), 023513.
    This paper describes advances in wearable fNIRS and high-density optical systems that allow cortical neuroimaging in more naturalistic conditions. It supports one of our central methodological points: we can move measurement closer to the lived world without confusing even a sophisticated hemodynamic signal with the totality of experience.

  10. Benitez, P.; Domeniconi, C. et al. (2023). “Análise da Viabilidade de Uso do fNIRS em Atividades Educacionais com Crianças e Jovens com Deficiência Intelectual e Autismo”. Revista Brasileira de Educação Especial, 29, e0158.
    This Brazilian study used fNIRS in a naturalistic clinical setting during educational activities involving children and young people with intellectual disability and autism. It provides a concrete Latin American example of how NIRS can measure cortical responses in contexts closer to real practices, helping move neuroscience beyond highly artificial laboratory tasks.

  11. BrainLatam (2026). “Bolhas vivas — Jiwasa, transcendência e o retorno ao planeta”.
    This article develops a distinction fundamental to the present text: “The screen transduces the world. The Body-Territory lives it.” It also presents 12D as a BrainLatam perceptual cartography, explicitly distinguishing it from claims about twelve physical dimensions or a neuroscientific consensus around exactly twelve senses.

  12. BrainLatam (2026). “O Logos se cristalizou em IA — por que ainda precisamos aprender a pensar?”.
    This article argues that generative AI contains only part of the linguistic and logical-formal production preserved in human archives and asks: “Which Logos crystallized?” The present Blog 01 moves one stage backward. If archives already contain absences, we must ask how particular events became records in the first place. Before the Logos crystallized into AI, the world had to crystallize into data.




#eegmicrostates #neurogliainteractions #eegmicrostates #eegnirsapplications #physiologyandbehavior #neurophilosophy #translationalneuroscience #bienestarwellnessbemestar #neuropolitics #sentienceconsciousness #metacognitionmindsetpremeditation #culturalneuroscience #agingmaturityinnocence #affectivecomputing #languageprocessing #humanking #fruición #wellbeing #neurophilosophy #neurorights #neuropolitics #neuroeconomics #neuromarketing #translationalneuroscience #religare #physiologyandbehavior #skill-implicit-learning #semiotics #encodingofwords #metacognitionmindsetpremeditation #affectivecomputing #meaning #semioticsofaction #mineraçãodedados #soberanianational #mercenáriosdamonetização
Author image

Jackson Cionek

New perspectives in translational control: from neurodegenerative diseases to glioblastoma | Brain States