Different Types of Data
Image copyright Adobe Stock
Quantitative and Qualitative
Many people first think of data as something like what’s in a table or spreadsheet of numbers and words. But data can take a variety of formats.
Quantitative Data – Involves a measurable quantity. Some examples are length, mass, temperature, and time.
Qualitative Data – Involves a descriptive judgment using concept words instead of numbers. Gender, country name, animal species, and emotional state are examples of qualitative information.
A graphic from UCLA libraries discussing the factors which shape what data is. Shared under a CC-BY-NC-SA license.
Depending on how you look at data, you might approach it quantitatively or qualitatively. It is important to have clear objective that helps you frame the information you have. For example, gender categories are usually discussed qualitatively, but in the context of conducting a demographic analysis, you might consider the number of people associated with each gender category.
All the examples below usually contain both qualitative and quantitative data.
This section is adapted from Choosing Sources, an open access resource from the Teaching and Learning team at Ohio State University. It is shared under a CC-BY 4.0 International License.
Governmental and Research Data
What are the main types of data you might encounter? The easiest to find are open data, that is, datasets collected or made available with access in mind. These are commonly of two types: governmental or research data. Governments and research institutions (which often overlap) sometimes have open data policies to promote transparency and facilitate information sharing. For example, Canada, Quebec, Montreal, and Concordia all have policies and platforms for collecting and sharing data in a systematic and predictable way that is accountable to their publics.
These databases reflect the activities those institutions want to emphasize and publicize, such as trade information, demographic information, or historical records. As such, they contain numerical information, such as municipal budgets, but also qualitative data, such as the results of a survey where people share their opinions.
These should be critically assessed, considering they are influenced by the politics of those institutions. For example, many countries had overtly legal forms of racial or gender discrimination into the second half of the 20th century. Census data dating prior to the enactment of anti-discrimination laws may use language or classifications of people that are offensive, racist, and do not reflect modern scientific consensus. Numerical data were often used to justify discriminatory policies as scientific or rational.
Business Data
A significant portion of data produced and used by private businesses is accessible to researchers and even the public. Regulations for which data businesses are required to share varies greatly between locales, but these data are aggregated and made accessible via a variety of databases and providers. Some data is freely available, while some is only available through proprietary (paywalled) databases. See, for example, Concordia Library’s Business, Company and Industry Data databases or our Industry Ratios and Reports page.
Medical Data
Medical data is both very important and very protected. In some cases, medical datasets are carefully scrubbed of personal information prior to being shared with the public for research purposes. This was not always the case: in the 20th century, many scientists who viewed some populations as less worthy of protection shared their information in ways considered disrespectful, if not illegal, today. See, for example, Henrietta Lacks’ legacy in medical history (Johns Hopkins Medicine, n.d.).
When collecting medical data with future sharing in mind, or when working with medical data, consent and future consequences are important to keep in mind.
Boundaries are Malleable
As you can probably begin to tell, boundaries between datasets are not rigid or well defined. Some medical datasets are made public by governments, some by corporations. Some business records are derived from tax information, which are also shared publicly by governments. As the boundaries between public and private are porous, so are the datasets they produce. These categories should be thought of as starting points which are put in motion by research and inquiry, not static delimitations that represent the truth. In the same sense, datasets are always assembled by and for people, with specific biases and agendas and intentions. In some cases, as in linguistics, even the boundary between qualitative and quantitative can shift: some words are read as instances of sounds or ideas, while in other cases they are for those sounds or ideas.
Clear context and objectives, as well as familiarity with relevant precedents, will allow you to engage with that information as effectively as possible.
What to do with data?
Accessing and downloading data is only the first step. Once you have the data you need, there is the arguably harder project of interpreting it critically and making that interpretation legible to your audience. In other words, you will probably be using data to produce new knowledge, whether it’s in articles, blog posts, on social media, in Youtube videos...
As explained in our Information Source Types unit, there are also many ways to classify, explore and access these information resources.
Resources
References
Johns Hopkins University School of Medicine. The Legacy of Henrietta Lacks no date.
Teaching and Learning, Ohio State University Libraries. Choosing & Using Sources: A Guide to Academic Research, Chapter 2. Quantitative or Qualitative. 2015.