Visually (Mis)representing Data

An image of a drawing of a rabbit's head / a duck's head which is drawn to look like either at the same time depending on how your brain chooses to identify it.

Duck or Rabbit? Public domain digital scan of a woodcut originally printed in the 23 October 1892 issue of Fliegende Blätter.

How do you see, read and use data every day?

Visualizing information is one of the most effective techniques used in communicating complex concepts, phenomena, and other scientific information. It enables some of us to see patterns, trends, and outliers in quantitative interpretations of our social and physical world. The last two decades, computational tools have made sophisticated graphing techniques more accessible. This has made visualization more common in all kinds of information resources, from academic journals to blogs, social media posts, and personal communication.

Because graphic representations of data are so common, visual data literacy is not a skill for experts: it is helpful for anyone who hopes to critically use data and graphs in their everyday life and academic or professional work.

In our other unit on Different Types of Data, we discuss the distinction between information and data. In our Quick Thing for Digital Learning on Data Visualization, we highlight the way that images representing data, for some of us, play a significant role in understanding complex concepts and ideas about the world that surrounds us.

This section adapts material from the Critical Data Literacy: Strategies to Effectively Implement and Evaluate Data Visualizations open resource published at Toronto Metropolitan University by Nora Mulvaney, Audrey Wubbenhorst and Amtoj Kaur. It is shared under a Creative Commons - Attribution 4.0 International License.​​

A brief history of seeing data

Representing information with graphs, charts and other clever visual or physical aids predates computers and digital information by centuries (Funkhouser, 1937). Contemporary data visualization has its roots in the nineteenth century, when the scientific practice of statistical data collection coincided with cheapening and diversifying printing techniques, which made including sophisticated graphics in books and magazines more accessible. This period saw the popularization of pie and bar charts, the histogram, line graphs, time-series plots, contour plots, scatterplots, etc. (Friendly, 2008a & b). Some famous and still relevant examples include the innovative work of Black social scientist W.E.B. DuBois, whose work visually quantifying the causes and consequences of racism remain relevant today (Mansky 2018). Beyond the continued relevance of discrimination, the techniques and political framing used by DuBois have influenced contemporary projects such as the Los Angeles Times' Data Desk mapping inequality across L.A. website (2024) or the Guardian's Data Editor Mona Chalabi's series of new graphics representing wealth gaps in the United States in the 2010's (2017) using DuBois' visual language.

A graph hand drawn and colored by W.E.B. DuBois c.1900. The line graph shows percentage of African Americans in slavery between 1790 and 1870.

Public domain scan of an original graph by W.E.B. DuBois, exhibited at the Paris World Fair in 1900.

Scan of a public domain from Wikimedia Commons: Example of polar area diagram by Florence Nightingale (1820–1910). This "Diagram of the causes of mortality in the army in the East" was published in Notes on Matters Affecting the Health, Efficiency, and Hospital Administration of the British Army and sent to Queen Victoria in 1858. This graphic indicates the annual rate of mortality per 1,000 in each month that occurred from preventable diseases (in blue), those that were the results of wounds (in red), and those due to other causes (in black). The legend reads: The Areas of the blue, red, & black wedges are each measured from the centre as the common vertex. The blue wedges measured from the centre of the circle represent area for area the deaths from Preventable or Mitigable Zymotic diseases, the red wedges measured from the centre the deaths from wounds, & the black wedges measured from the centre the deaths from all other causes. The black line across the red triangle in Nov. 1854 marks the boundary of the deaths from all other causes during the month. In October 1854, & April 1855, the black area coincides with the red, in January & February 1856, the blue coincides with the black. The entire areas may be compared by following the blue, the red, & the black lines enclosing them.

Public domain scan of a graph from 1858 by Nightingale and her team.

Florence Nightingale (1820-1910) was a nurse and social scientist who was employed by the British government to help with war casualties. Over the course of her career, she collaborated with doctors, statisticians and artists to develop innovative visual tools which supported her policy recommendations. For example, she realized many more soldiers died after battles because of poor care for injuries than on the battlefield itself. This led to graphs such as the one above, a "polar area chart" visualizing causes of death in relevant conflicts in a way which was more legible and immediate to the people reading her reports (Hedley, 2020).

By the time graphing software became more and more popular in the 1980's and 1990's, decades of visual idioms, scientific subcultures and graphic design had happened, foreshadowing the interactive, web-based and/or big data approaches common today.

Graphs can also be misleading...

For all of their illuminating potential, graphs can also be designed to subtly (or not-so-subtly) manipulate facts and people's opinions. Because graphs tend to visually organize data in a formal, well-defined way, most misrepresentations of data work via warping or omitting part of the scale or the data. This results in unclear, incomplete, and misleading graphs. Types of misleading data distortions also include using an inappropriate type of graph (e.g. a pie chart when a bar graph would be clearer), implying correlations between trends indicate causation, unnecessary visual effects (such as shading or perspective), or various kinds of color manipulation... Just as there are endless clever ways to use graphs, there are also many ways to misuse graphs.

A television image with Blue Jays pitcher R.A. Dickey standing on right and a bar graph depicting his throw speed on the left. The left bar shows 77.3 mph for 2012 but the bar is twice as high as the 2013 bar, labeled 75.3mph. The two miles per hour difference is only 2.6 % of 77.3mph and the bars should be almost exactly the same height.

A misleading graph dramatizing a baseball pitcher's lower throw velocity between two seasons. Although the difference in speed is only 2 mph (2.6% of the initial value) the first bar is more than twice as high as the second bar in this bar graph.

... or just unfortunately designed.

Even when there is no ill intent, the ubiquitous presence of graphs in scientific knowledge has resulted in the publication of graphs that are confusing only because they are too dense, not labeled clearly, or otherwise haphazardly designed (Wainer, 1984). In some cases, there can also be real, intentional compromising of legibility for the sake of density. The extent to which scientific publications should be legible to non-experts remains a topic of debate. Regardless of the reason, the aesthetics of difficult to read scientific graphs and technical drawings have become a popular trope of their own. See, for example, the "ugly charts" page compiled by FlowingData.com.

A very dense, multicolored graph overlaying multiple kinds of plots with different scales and values. The scientists used a variety of measurement techniques and criteria to classify planetary bodies, this graph shows all the results simultaneously for a focused region from the previous graph in the paper.

A very dense, multicolored graph overlaying multiple kinds of plots with different scales and values. The scientists used a variety of measurement techniques and criterias to classify planetary bodies, this graph shows all the results simultaneously for a focused region from the previous graph in the paper. From Zeng et al. 2019.

Critically reading graphs

What are some strategies to critically read data visuals?

  1. Stop and Slow Down: as when evaluating any information source, the first step in critically reading a graph is to stop and consider the circumstances that have led to this particular image. Was it recommended algorithmically? Did someone send it to you? Is it part of an article you are reading? In every case, ask yourself: who put this there, and why? 
  2.  Separate the Scaffolding from the Visual Encoding: the best way to understand what a graph does is to separate the scaffolding (title, legends, scales, annotations, sources, bylines, etc.) from the visual encoding (the symbols or visual elements representing data).
  3. Focus on Scaffolding First: consider if the type of graph enables clear comparisons of different entries. This will give you some background about the choices made by the authors of the graph.
  4. Consider Content and Visual Encoding Second: what is the visual language used by the graph? what size, colour, and groupings are present? These are the rules of organization of data in this graph. The more complex the rules, the harder it becomes to decipher and the easier it is for information to be manipulated.
  5. Spot Patterns and Relationships: now that you have a better sense of the rules and contents of this graph, reflect on the patterns it suggests and the relationships it highlights. Do these fit the topic, and perspectives leading up to the graph? Do you notice any unacknowledged bias or agenda?
  6. Examine Your Own Biases: if the graph doesn't have noticeable biases, it is likely that it is because those present simply align with your own biases and expectations. Confirmation bias is the very real phenomena of reading what you believe in a source regardless of its actual content. Are you reading to learn, or to find facts you already know?

More generally, once you've identified where in an argument or project a graph fits, and what your own position relative to that argument is, you will gain a more reflexive and critical understanding of visual data representations.

This section adapts material from the Critical Data Literacy: Strategies to Effectively Implement and Evaluate Data Visualizations open resource published at the Toronto Metropolitan University by Nora Mulvaney, Audrey Wubbenhorst and Amtoj Kaur. It is shared under a Creative Commons - Attribution 4.0 International License.

Resources

References

Friendly, Michael. The Golden Age of Statistical Graphics. Statistical Science 23 (4): 502–35. 2008a.

Friendly, Michael. A Brief History of Data Visualization. In the "Handbook of Data Visualization" edited by Chun-houh Chen, Wolfgang Härdle, and Antony Unwin. Springer. pp15-56. 2008b.

Funkhouser, H. Gray. Historical Development of the Graphical Representation of Statistical Data. Osiris 3 (January):269–404. 1937.

Hedley, Alison. Florence Nightingale and Victorian Data Visualisation. Significance 17 (2): 26–30. 2020.

Mansky, Jackie. W.E.B. Du Bois’ Visionary Infographics Come Together for the First Time in Full Color. Smithsonian Magazine. November 15, 2018.

Wainer, Howard. How to Display Data Badly. The American Statistician 38 (2): 137–47. 1984.

Zeng, Li, Stein B. Jacobsen, Dimitar D. Sasselov, Michail I. Petaev, Andrew Vanderburg, Mercedes Lopez-Morales, Juan Perez-Mercader, et al. Growth Model Interpretation of Planet Size Distribution. Proceedings of the National Academy of Sciences 116 (20): 9723–28. 2019.