Heat to Height: Comparing the Physical Dimensionality of Heat Maps

Ph.D. Final Examination

Tyler Wiederich

April 17, 2026

Outline

  • Motivation
  • Study Design
  • Results
  • Conclusions
  • Contributions and Future Work

Motivation

Motivation

Data is everywhere, but how do we communicate it through visualizations? We design graphics to account for:

  • Audience
  • Clarity
  • Accessibility

Many types of charts exist for displaying different types of data.

What alternative methods can we use to better display it?

Heat maps

Heat maps are a type of data visualization that convey data in three dimensions.

  • Two dimensions represented by spatial coordinates

  • Other dimension (typically a continuous response variable) represented by alternative visual encoding

    • color/fill, shading, height

Data Physicalization

Many advancements in technology have allowed for greater artistic creativity in displaying data.

One of the earliest 3D-printed data visualizations. Source: Wall Street Journal

3D-printed charts

What are the potential benefits of using 3D-printed charts for statistical graphics?

  • Increased interaction
  • Less reliance on color
  • Engagement and attention
  • Better accessibility

Previous Research

To date, very few empirical studies have been conducted using physical 3D charts.

Many instances of 3D-printed charts are created for artistic representations!

Examples

Jansen, Dragicevic, and Fekete (2013)

Blog post from VisWorld

Tree map (Huron et al. 2023)

Wall Street Journal

Examples of 3D-printed charts

Overview of Our Research

Previous research suggests that 3D charts may perform better when all three dimensions are used for displaying information (Barfield and Robless 1989; Fisher, Dempsey, and Marousky 1997; Kraus et al. 2020).

Does this hold for numerical estimations from physical 3D heat maps compared to 2D static heat maps and interactive digital 3D heat maps?

There are many possible conditions for which charts can be rendered that could influence numerical estimations.

To what extent do underlying datasets and comparison types affect numerical estimations from these charts?

Methods

Stimuli

The design of the 3D heat map experiment uses the method of constant stimuli: ratios are estimated with respect to one stimuli height that remains the same.

\[ S=\text{Stimuli} \]

Setting 50 as the constant and 90 as the maximum, a sequence of stimuli are chosen by equally partitioning the ratios between \(50/50=1\) and \(50/90\approx0.556\). The same ratios are used when setting 50 as the maximum in the stimuli pair.

Stimuli Values

Heat Map Data

We used two datasets created using a mixture distribution for equations for spheres and random noise as a reference distribution. Final values were scaled between 0 and 100.

  • Set 1: \(f_1(X,Y)=\sqrt{7^2-(X-\bar{X})^2-(Y-\bar{Y})^2}\)
  • Set 2: \(f_2(X,Y)=\sqrt{7^2+(X-\bar{X})^2+(Y-\bar{Y})^2}\)

Set 1

Set 2

Stimuli Placement

We strategically placed stimuli so that they looked natural with respect to the simulated dataset.

Placement of stimuli on Set 1

Placement of stimuli on Set 2

Chart Types

(a) 2D Digital
(b) 3D Digital
(c) 3D Printed
Figure 1: Chart types representing heat map Set 1.

Experimental Design

For a full replicate, there are \(3\times2\times9=54\) treatment combinations:

  • 3 media types (2dd, 3dd, 3dp)
  • 2 datasets
  • 9 pairs of stimuli

Way too many trials for a single participant’s attention span!

Our main interest is the difference between media types, measured at a given ratio and dataset. To accomplish this and to reduce the number of trials per participant, we use 4 of the 9 possible stimuli pairs to create blocks.

\[ 2\times3\times4=24 \]

Responses

For each trial in the experiment, we ask two question adapted from Cleveland and McGill (1984).

  1. Which value in a stimuli pair represents a larger quantity?
  2. If the larger value in the stimuli pair represents 100 units, how many units is the smaller value?

Question 1 has three options: one for each value in the pair and one option for if the values are the same.

Question 2 has a slider ranging from 0 to 100, and the initial position is randomly positioned

Shiny application

Participant Recruitment

We recruited our subjects by integrating an experiential learning project into the STAT 218 curriculum at UNL. The data collection period was from Summer 2025 to Fall 2025.

For data to be collected, students had to meet the following criteria:

  1. Be at least 19 years of age
  2. Consent to participation

Students were instructed to complete the experiment at communal office hour locations.

  • Online sections were exempt from this, and 3D-printed charts were removed from their trials.

Results

Demographics

259 students completed the entirety of the experiment. Of these, 92 students were enrolled in online sections of the course.

Completion Times

Figure 2: Histogram of experiment completion times. Online participants had generally shorter completion times due to a decreased number of trials.

Careless Respondents

Figure 3: Three potential estimation strategies from participants. Many participants followed instructions to estimate ratios. However, some participants appeared to estimate the difference between stimuli pairs or submitted random values.

How to deal with careless respondents?

While careless respondents are an acknowledged issue, there is no general consensus on how to handle them.

  • Preventative actions (attention checks, etc.)
  • Post-hoc analysis (removal, mixture models, etc.)

Our first question was actually designed as a quasi-attention check! For each participant, we calculate p-values using a one-sided Binomial test:

\[ H_0: \pi\leq2/3\text{ vs. }H_1:\pi>2/3 \]

where \(\pi\) is the probability of correctly guessing the larger value.

Was this participant answering Q1 correctly at least twice as often as incorrect options?

Figure 4: Proportion of correct responses to Question 1 with all participants

Figure 5: Proportion of correct responses to Question 1 without random guessers

Analyzing Error

We calculate error as follows:

\[ y=\log_2(|\text{User Response}-100\times(\text{True ratio})|+1/8) \]

We excluded:

  • suspected random guessers
  • stimuli pair where values were the same
  • estimates where Q1 were incorrect

Various models produced the same conclusions!

Linear Mixed Model

\[ Y_{ijklm}=\mu+S_i\times M_j\times P_k+\gamma_{lm}+\omega_{ijlm}+\epsilon_{ijklm} \]

where

  • \(Y=\log_2(|\text{User Response}-100\times(\text{True ratio})|+1/8)\)
  • \(\mu\) is the intercept
  • \(S_i\times M_j \times P_k\) is the fixed effects (and interactions) of Set, Media type, and stimuli Pair
  • \(\gamma_{lm}\) is the random effect of participant and block
  • \(\omega_{ijklm}\) is the random effect of media and set for each participant
  • \(\epsilon_{ijklm}\) is the random error
  • All random effects i.i.d. \(N(0,\sigma^2)\) with separate \(\sigma^2\) and are independent
ANOVA Table
F Df Df.res Pr(>F)
(Intercept) 332.257 1 2477.652 0.000
set 0.114 1 2745.681 0.735
media 3.189 2 2757.181 0.041
pair_id 1.919 7 2626.034 0.063
set:media 0.401 2 2747.072 0.669
set:pair_id 1.856 7 2681.555 0.073
media:pair_id 0.424 14 2661.569 0.968
set:media:pair_id 0.641 14 2669.463 0.833
Figure 6: Distributions of log error overlaid with estimated marginal means and their 95% confidence intervals. Differences were only detected between the 2D heat map and both types of 3D heat maps, where the 2D heat map had evidence of larger errors.
Figure 7: Estimated marginal means for media types and their 95% confidence intervals.

Figure 8: Estimated marginal means of stimuli pair and dataset. Bars sharing the same letter are not significantly different from one another.

Time per trial

(a) Frequency plot
(b) Time vs. Error
Figure 9: Visual summary of time spent on each trial. 137 trials are omitted due to lasting longer than 60 seconds.

Interactions with digital 3D heat map

(a) Histogram
(b) WebGL clicks vs. Error
Figure 10: Participant interactions with the digital 3D heat maps.

Interactions with the slider

(a) Histogram
(b) Slider clicks vs. Error
Figure 11: Participant interactions with the slider.

Initial Placement of Slider

For each trial, the slider was randomly positioned between 0 and 100.

Figure 12: Effect of initial slider position on the submitted slider value.

Conclusions and Future Work

Main findings

We note the following results:

  • Digital and physical 3D heat maps produced lower error than 2D heat maps
  • No detectable differences between digital and physical 3D heat maps
  • Differences attributed to dataset and stimuli pair

Areas for future work

As one of the first studies to empirically evaluate physical 3D charts, there are many areas left to explore.

  • Colors and fills of 2D and 3D heat maps
  • Just noticeable differences
  • Types of estimations (\(A/B\), \(A/(A+B)\), \(B-A\), etc.)
  • Development of tools to produce 3D charts with considerations for 3D printing

Contributions to literature

Contributions

Study 1: Partial replication of Cleveland and McGill (1984) to compare digital and physical representations of bar charts.

No significant differences in ratio estimations detected from type of chart (\(n=38\))

Contributions

Study 2: Students completed a series of reflections regarding their participation in the experiment.

What information do students learn from participating in an experiment, both as participants and consumers of scientific knowledge?

Many students connected ideas from the course material! (e.g., variable identification and generalizability)

Contributions

Study 3: taking inspiration from Cleveland and McGill (1984) and Study 1, we extend our scope to use heat maps.

When all dimensions are used for conveying information, 3D heat maps produced lower error rates for estimations of ratios than 2D heat maps.

No detectable differences between digital and physical 3D heat maps.

Questions?

Appendix

Multiple Models

ANOVA Table p-values for model terms
Term All participants, all responses All participants, Q1 correct Filtered participants, all responses Filtered participants, Q1 correct
set 0.340 0.705 0.939 0.735
media 0.010 0.018 0.035 0.041
pair_id 0.244 0.333 0.089 0.063
set:media 0.297 0.711 0.543 0.669
set:pair_id 0.003 0.145 0.001 0.073
media:pair_id 0.815 0.740 0.982 0.968
set:media:pair_id 0.489 0.764 0.255 0.833

Figure 13: Interaction plots for media type and response filtering, facetted by stimuli pair and dataset.

Figure 14: Interaction plots for dataset and response filtering, facetted by stimuli pair and media type.

Figure 15: Interaction plots for stimuli pairs and response filtering, facetted by dataset and media type.

Excluding Stimuli Pair 5

Figure 16: Counts of response behavior for Stimuli Pair 5. This pair had identical values, which means that true solutions indicates marking that they were the same value and positioning the slider at 100.

Generalized Additive Model

\[ y=\mu+S_i\times M_j + s_i^S(R)+s_j^M(R)+P_m+\epsilon \qquad(1)\]

where

  • \(y=\log_2(|\text{Error}|+1/8)\)

  • \(S_i\times M_j\) is the fixed effects for main effect and two-way interactions of dataset \(S_i\) and media type \(M_j\)

  • \(S_i^S(R)\) is a thin-plate spline function for ratio accounting for dataset

  • \(S_j^M(R)\) is a thin-plate spline function for ratio accounting for media type

  • \(P_m\) is the random effect of the \(m^{th}\) participant

  • \(\epsilon\) is random error

GAM (All Participants)


Family: gaussian 
Link function: identity 

Formula:
q2_error_cm ~ set * media + s(target_ratio, k = 4, by = media) + 
    s(target_ratio, k = 4, by = set) + s(user_id, bs = "re")

Parametric Terms:
          df      F  p-value
set        1  0.030    0.863
media      2 12.190 5.29e-06
set:media  2  0.211    0.810

Approximate significance of smooth terms:
                             edf  Ref.df     F p-value
s(target_ratio):media2dd   1.002   1.003 0.058   0.813
s(target_ratio):media3dd   1.972   2.350 1.476   0.200
s(target_ratio):media3dp   1.047   1.428 1.612   0.268
s(target_ratio):setset1    1.005   1.009 0.650   0.422
s(target_ratio):setset2    1.022   1.042 0.165   0.719
s(user_id)               217.826 258.000 6.423  <2e-16

GAM (Filtered Participants)


Family: gaussian 
Link function: identity 

Formula:
q2_error_cm ~ set * media + s(target_ratio, k = 4, by = media) + 
    s(target_ratio, k = 4, by = set) + s(user_id, bs = "re")

Parametric Terms:
          df      F  p-value
set        1  0.279    0.598
media      2 13.634 1.28e-06
set:media  2  0.450    0.638

Approximate significance of smooth terms:
                               edf    Ref.df     F  p-value
s(target_ratio):media2dd 3.368e-04 6.524e-04 0.186 0.991204
s(target_ratio):media3dd 1.006e+00 1.011e+00 2.510 0.111735
s(target_ratio):media3dp 1.011e+00 1.022e+00 0.791 0.369414
s(target_ratio):setset1  2.232e+00 2.599e+00 7.748 0.000327
s(target_ratio):setset2  2.349e+00 2.702e+00 6.397 0.000404
s(user_id)               1.396e+02 1.600e+02 7.105  < 2e-16

References

Barfield, Woodrow, and Robert Robless. 1989. “The Effects of Two- or Three-Dimensional Graphics on the Problem-Solving Performance of Experienced and Novice Decision Makers.” Behaviour & Information Technology 8 (5): 369–85. https://doi.org/10.1080/01449298908914567.
Cleveland, William S., and Robert McGill. 1984. “Graphical Perception: Theory, Experimentation, and Application to the Development of Graphical Methods.” Journal of the American Statistical Association 79 (387): 531–54. https://doi.org/10.1080/01621459.1984.10478080.
Fisher, Samuel H., John V. Dempsey, and Robert T. Marousky. 1997. “Data Visualization: Preference and Use of Two-Dimensional and Three-Dimensional Graphs.” Social Science Computer Review 15 (3): 256–63. https://doi.org/10.1177/089443939701500303.
Huron, Samuel, Till Nagel, Lora Oehlberg, and Wesley Willett, eds. 2023. Making with Data: Physical Design and Craft in a Data- Driven World. First edition. AK Peters Visualization Series. Boca Raton: AK Peters : CRC Press.
Jansen, Yvonne, Pierre Dragicevic, and Jean-Daniel Fekete. 2013. “Evaluating the Efficiency of Physical Visualizations.” In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, 2593–2602. Paris France: ACM. https://doi.org/10.1145/2470654.2481359.
Kraus, Matthias, Katrin Angerbauer, Juri Buchmüller, Daniel Schweitzer, Daniel A. Keim, Michael Sedlmair, and Johannes Fuchs. 2020. “CHI ’20: CHI Conference on Human Factors in Computing Systems.” In, 1–14. Honolulu HI USA: ACM. https://doi.org/10.1145/3313831.3376675.
Stusak, Simon, Moritz Hobe, and Andreas Butz. 2016. “If Your Mind Can Grasp It, Your Hands Will Help.” In Proceedings of the TEI ’16: Tenth International Conference on Tangible, Embedded, and Embodied Interaction, 92–99. Eindhoven Netherlands: ACM. https://doi.org/10.1145/2839462.2839476.
Stusak, Simon, Jeannette Schwarz, and Andreas Butz. 2015. “Evaluating the Memorability of Physical Visualizations.” In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, 3247–50. Seoul Republic of Korea: ACM. https://doi.org/10.1145/2702123.2702248.