Towards Personalized Annotation of Webpages for Efficient Screen-Reader Interaction
Full-text Seed edition of the ACM HT 2020 paper.

Authors: Hae-Na Lee, Vikas Ashok

Abstract

To interact with webpages, people who are blind use special-purpose assistive technology, namely screen readers that enable them to serially navigate and listen to the content using keyboard shortcuts. Although screen readers support a multitude of shortcuts for navigating over a variety of HTML tags, it has been observed that blind users typically rely on only a fraction of these shortcuts according to their personal preferences and knowledge. Thus, a mismatch between a user’s repertoire of shortcuts and a webpage markup can significantly increase browsing effort even for simple everyday web tasks. Also, inconsistent usage of ARIA coupled with the increased adoption of styling and semantic HTML tags (e.g., <div>, <span>) for which there is limited screen-reader support, further

make interaction arduous and frustrating for blind users.

To address these issues, in this work, we explore personalized annotation of webpages that enables blind users to efficiently navigate webpages using their preferred shortcuts. Specifically, our approach automatically injects personalized ‘annotation’ nodes into the existing HTML DOM such that blind users can quickly access certain semantically-meaningful segments (e.g., menu, search results, filter options, calendar widget, etc.) on the page, using their preferred screen-reader shortcuts. Using real shortcut profiles collected from 5 blind screen-reader users doing representative web tasks, we observed that with personalized annotation, the interaction effort can be potentially reduced by as much as 48 (average) shortcut presses.

CCS CONCEPTS


    Human-centered computing → Accessibility technologies;

    User studies.

KEYWORDS

Web screen-reading, transcoding, web accessibility.

ACM Reference Format: Hae-Na Lee and Vikas Ashok. 2020. Towards Personalized Annotation of Webpages for Efficient Screen-Reader Interaction. In Proceedings of the 31st ACM Conference on Hypertext and Social Media (HT ’20), July 13–15, 2020, Virtual Event, USA. ACM, New York, NY, USA, 6 pages. https://doi.org/10. 1145/3372923.3404815

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. HT ’20, July 13–15, 2020, Virtual Event, USA © 2020 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-7098-1/20/07...$15.00 https://doi.org/10.1145/3372923.3404815

1 INTRODUCTION

The Web has made in-roads into almost every aspect of modern society, as people are increasingly carrying out their everyday activities online such as shopping, social interaction, communication, news reading, flight and hotel reservation, etc. Websites are becoming more sophisticated with visually rich and dense content layouts in order to attract more traffic, and hence business. Unfortunately, these rich visual enhancements have made web interaction challenging and arduous for people who are blind and visually impaired (BVI) relying on non-visual screen-reader assistive technology such as JAWS, NVDA, VoiceOver, etc. [1, 18, 28], to interact with webpages.

A screen reader, as the name suggests, serially narrates screen content, and also enables the BVI to navigate the webpage content using special keyboard shortcuts (e.g., ‘H’ for the next heading). Although screen readers offer plenty of shortcuts to navigate a wide range of HTML elements, it is challenging and often impractical to remember all of the shortcuts (e.g., 100+ shortcuts in JAWS), and therefore blind users typically rely on only a handful of shortcuts based on their preference and knowledge [5, 13]. This limited lexicon of screen-reader shortcuts can significantly increase the navigation effort in webpages for the BVI, even more so in modern content-dense webpages laden with plethora of visually-appealing styling and semantic HTML elements. Also, inconsistent usage of accessibility-enhancing WAI-ARIA [27] attributes by web developers further exacerbates the interaction burden of the BVI.

In this paper, we explore the potential of personalized annotation of webpages, with an intent of ‘transparently’ improving efficiency of content access and navigation for screen-reader users. To the best of our knowledge, this work represents a seminal effort in personalized annotation for web accessibility, while other extant approaches [3, 7, 14] mostly propose ‘one-size-fits-all’ solutions aimed primarily at exposing the visually-encoded page semantics (e.g., structure, dynamic changes, etc.) to the BVI, with little to no focus on efficiency of access.

While annotating the entire webpage DOM to match the shortcut profile of a BVI user is extremely challenging and probably infeasible given the DOM sizes of modern webpages, in this preliminary work, we explore personalization methods to annotate only specific portions of a webpage DOM that correspond to certain semantically-meaningful entities such as menus, search results, filter options, sort options, calendar widget, forms, news articles, etc. Specifically, we built a browser extension PAN that leverages state-of-the-art information extraction techniques to identify the DOM subtrees corresponding to the different semantic entities, and then injects ‘dummy’ HTML nodes so as to facilitate quick access to these entities using only the shortcuts in the user’s personal

shortcut profile. PAN also maintains and updates the user’s shortcut profile; therefore every personalized transformation is ‘fresh’, reflecting the user’s most recent repertoire of shortcuts.

2 RELATED WORK

There has been significant research advancement in understanding and improving the accessibility of webpages [5, 8, 13, 19, 22, 24]. Web Automation and Assistants. Web automation techniques

[10, 12, 20, 23] enable the BVI to automate certain tasks, thereby facilitating quick execution of these tasks. In almost all of these techniques, the task scripts or macros for automation are created either by handcrafting [12, 23], or by user demonstration [10, 21, 26]. While these techniques indeed reduce interaction effort for the BVI, they are mostly useful only for repetitive tasks. Furthermore, end users have to not only spend considerable time and effort creating a script for each task, but also maintain and update scripts.

Accessibility assistants [4, 5, 11, 16] enable BVI users to use other input modalities, e.g., speech, audio-haptic input devices, etc., to quickly navigate webpage content. For example, Billah et al. [11] propose using a Dial input device to hierarchically navigate page content using rotate and press gestures. Gadde et al. [16], on the other hand, propose a speech-based interface that enables the BVI to get a quick overview of what is present on the current webpage and navigate to key sections using a few simple voice commands.

A common denominator underlying both automation and assistantbased approaches is that the BVI have to learn how to use new interfaces in addition to remembering screen-reader shortcuts and re-devise their browsing strategies, which can consume significant time and effort. Also, since many of these highly usable approaches (e.g., [5, 16]) require tighter integration with third-party screenreader framework, their scope is limited to a few open-source screen readers. In contrast, our approach is transparent to the users. Therefore they can enjoy the benefits of faster content access and navigation while continuing to use their preferred screen-reader shortcuts. Annotating Webpages for Accessibility. The idea of annotating webpages (or transcoding) to promote accessibility has been explored in many previous works [3, 6, 7, 9, 14, 17, 29]. In an early approach, Asakawa et al. [3] propose injecting annotations (in the form of dummy HTML elements) into the existing webpages to convey the visual structural information of the page content to screen-reader users. However, since their approach requires human annotators to use a custom annotation vocabulary and create annotation files for each website, it is less likely to be adopted by web developers due to considerable overhead [7]. Instead of annotating the base document, Bechhofer et al. [7] propose using a predefined ontology to annotate the existing CSS stylesheets that can then be exploited by other transcoding tools to manipulate page content for improving accessibility.

Apart from annotating semantics, researchers have also explored injecting JavaScript and ARIA into the webpage DOM to improve accessibility [6, 14, 17, 29]. For instance, Brown et al. [14] propose a method to inject JavaScript into webpages so as to capture and classify dynamic changes on these pages and then notify the details of these changes to the screen-reader users via an ARIA live region. On the other hand, a recent work [6] uses visual saliency deep neural networks to identify the important parts or ‘hot-spots’ of

the webpage, and then automatically inject ARIA landmark roles into their corresponding branches in the page DOM.

While all the aforementioned approaches help improve content access for the BVI, they are primarily concerned with explicitly exposing the visually-encoded semantics, with little focus on efficiency of access. Moreover, as in case of web automation, most of them require users to learn new interaction techniques; none of them take into account the screen-reader expertise or the personal shortcut profile of end users. Therefore, in this paper, we try to fill this gap by exploring personalized and transparent transformation techniques for the BVI that not only enable quick access to the various semantic structures in the page, but also do not require BVI users to change or learn new interaction behavior.

3 PERSONALIZED ANNOTATION FOR EFFICIENT WEB SCREEN READING

We built a browser extension PAN (for the Chrome browser) that monitors a user’s shortcut presses to build and maintain a corresponding profile. Whenever a new page is loaded in the browser, the extension leverages the most recent profile information to accordingly annotate the webpage. As tailoring the annotation for each HTML element in the DOM is expensive given the large size of modern sophisticated webpages, and often also unnecessary and counterproductive (e.g., the users would rather skip irrelevant content such as advertisements), the extension only focuses on annotating a few generic semantically-meaningful segments (e.g., search results, webpage menus, calendar, filter options, etc.). Once these segments are identified in the DOM, the extension then uses one of the personalized techniques (see Section 3.3) to inject new ‘dummy’ HTML nodes adjacent to the root nodes of these segments. The details are provided next.

3.1 Identifying Semantically-Meaningful Segments in a Webpage

Identifying the semantically-meaningful segments in a webpage corresponds to locating and labeling the corresponding subtrees in the overall page DOM (see Figure 1a). Automatically identifying every semantic entity on any arbitrary webpage is impractical, given the vast number of different types of entities (including custom widgets) being used all over the web. Therefore, PAN only focuses on identifying specific Generic Segments (GSs) that are very commonly used across multiple webpages, and also tend to be the key segments in these pages. In this paper, we considered 8 such Generic Segments – list of items such as products, flights, inbox, search results; news articles; menus; list of filter options; forms; calendar widget; list of sort options; and sidebars.

To locate and label the subtrees corresponding to the various GSs on the page, the extension leverages a wide range of existing techniques in the related literature [5, 15, 22, 30–32]. Each of these existing approaches focuses on identification of specific GSs. For example, Álvarez et al. [32] present a technique to identify and extract the list of items such as search results, by taking advantage of repetitive patterns and other similarities such as XPath and visual formatting. On the other hand, the machine-learning based technique in [22] is capable of monitoring changes in the DOM and identifying widgets such as calendars, suggestions lists,

<body>

Col1

Col2

Col3

Col4

Col5

Root of GS

Col7

Col8

• • •
Root

• • •
Root

• • •
Root

• • •
Root

• • •
Root

Root

Root

Root

























Basic shortcut

Col2

Specific shortcut


(a) (b)

the List of Items GS. The arrows illustrate serial screen-reader navigation over the DOM elements. The personalized annotator uses the shortcut profile to inject a dummy HTML element <h1> with "List of items" as its innerHTML text.

web chats, etc., using several custom-developed features (including ARIA markups) extracted from the DOM. Similarly, [2, 5, 25] present extraction techniques to identify menus, news-article contents, and sort options respectively. Identifying forms is simple and straightforward by given that we just need to search for the <form> HTML tags in the DOM.

We evaluated the identification algorithms on a custom-built test dataset consisting of 948 segments from 256 webpages belonging to 100 popular websites covering various domains such as shopping, news, weather, email, social media, flight reservation, government, etc. The average accuracies of identification for the 8 GSs are shown

Table 1: Identification accuracy on test dataset comprising 948 segments from 256 webpages.

in Table 1. As seen in Table 1, the algorithms yielded a high average accuracy, i.e., correctly identified the GS subtrees. Also, inaccurate identification results could be categorized into three types: (i) partial identification, (ii) extraneous identification; and (iii) failed identification. There were no false positives, i.e., identifying one generic entity incorrectly as some other entity.

3.2 Shortcut Navigation and Profile

Figure 1b shows the linear screen-reader navigation of the HTML DOM. When a page is loaded in the browser, by default, the screenreader focus is at the top of the page corresponding to the <body> element in the DOM. Depending on the shortcut pressed, the screenreader focus moves to the next corresponding element, and the text

Identification Accuracy (%) Segment Exact Partial Extra Undetected List of Items 84.6 8.6 3.7 3.1 News Article 92.2 6.6 0 1.2 Webpage Menu 74.4 12.5 9.4 3.7 Filter Options 78.5 17.4 0 4.1 Forms 100 0 0 0 Calendar 95.6 0 0 4.4 Sort Options 87.7 0 5.4 6.9 Sidebars 68.3 12.2 2.8 16.7

Table 1: Identification accuracy on test dataset comprising 948 segments from 256 webpages.

associated with that element is narrated to the user. For example, assuming the use of JAWS screen reader, when the user presses an 𝐻 shortcut, the focus moves to the next heading and the text within the heading is read out to the user. Users can also navigate backwards (e.g., previous heading) using similar shortcuts (e.g., 𝑆𝐻𝐼𝐹𝑇 + 𝐻 ).

PAN maintains the user’s shortcut profile in the form of a set 𝑆 of pairs < 𝑠, 𝑓 (𝑠) >, where 𝑠 is a screen-reader shortcut and 𝑓 (𝑠) is the corresponding frequency of usage. The shortcuts supported by screen readers can be categorized into two types: (i) Basic shortcuts - shortcuts such as arrow keys that enable users to navigate the HTML DOM element-by-element regardless of their type (i.e., tagname); and (ii) Specific shortcuts - shortcuts such as 𝐻, 𝑃, 𝐿, etc., which are exclusive for specific HTML element types. As observable in Figure 1b, Specific shortcuts enable users to skip more content and reach the desired content much faster compared to that with type-agnostic Basic shortcuts; therefore the extension only considers Specific shortcuts in a profile for making personalized annotations. Furthermore, the extension continuously updates the profile after each shortcut press during web browsing so as to keep abreast of changes in user’s shortcut knowledge and preferences, thereby ensuring that ‘fresh’ annotations are made to a page whenever it is loaded in the browser.

3.3 Personalized Annotation

The goal of annotation is to reduce the time and the number of shortcuts needed to navigate to the beginning of the GSs in a webpage. Personalized annotation achieves this objective by injecting ‘dummy’ HTML nodes (adjacent to the root of GSs) that a user

is highly likely to reach quickly, given their shortcut profile. For example, as illustrated in Figure 1b, if the user is familiar with, and frequently presses the heading ‘H’ shortcut, but is not familiar with the list ‘L’ shortcut, injecting a ‘dummy’ heading with the text “List of Items” before the list increases the chances of the user navigating to the search results faster than without any annotation. This way, by creating new paths that are tailored to match the user’s observed shortcut knowledge and preferences, personalized annotation improves efficiency and ease of content access.

Identification Accuracy (%)

Col2

Col3

Col4

Col5

Segment

Exact

Partial

Extra

Undetected

List of Items

84.6

8.6

3.7

3.1

News Article

92.2

6.6

0

1.2

Webpage Menu

74.4

12.5

9.4

3.7

Filter Options

78.5

17.4

0

4.1

Forms

100

0

0

0

Calendar

95.6

0

0

4.4

Sort Options

87.7

0

5.4

6.9

Sidebars

68.3

12.2

2.8

16.7

While many schemes can be potentially devised to determine how to inject personalized annotations for GSs, in this preliminary work, we focus on the following custom-designed approaches:

    Path-Agnostic: For each GS, PAN injects an HTML element

corresponding to the most frequently used Specific shortcut in the profile. Note however, under this scheme, PAN will not inject a dummy node if the root node of the GS, by default matches the most-used Specific shortcut. Specifically, if 𝐺 is the set of GSs on the page; 𝐿 is the mapping from Specific shortcuts to the corresponding HTML elements, and 𝑆 = < 𝑠, 𝑓 (𝑠) > is the shortcut profile, then the annotation node 𝑦 injected is given by:

∀𝑔 ∈ 𝐺, 𝑦 (𝑔) = 𝐿(𝑥), where 𝑥 = argmax

𝑠 ∈𝑆

𝑓 (𝑠)

ID

Shortcut Profile

P1

<‘E’, 136>, <‘Tab’, 112>, <‘P’, 85>, <‘Shift+E’, 45>, <‘A’,
14>, <‘Shift+A’, 6>, <‘G’, 2>

P2

<‘B’, 65>, <‘F’, 47>, <‘Shift+B’, 34>, <‘T’, 31>, <‘M’, 4>

P3

<‘Tab’, 489>, <‘Shift+Tab’, 325>, <‘H’, 183>, <‘Shift+H’,
96>, <‘T’, 43>, <‘P’, 18>, <‘Shift+T’, 4>, <‘Z’, 1>

P4

<‘H’, 326>, <‘Shift+H’, 274>, <‘Tab’, 257>, <‘B’, 173>,
<‘Shift+Tab’, 164>, <‘Z’, 102>, <‘T’, 53>, <‘Shift+Z’, 39>,
<‘F’, 22>, <‘Shift+B’, 7>, <‘C’, 2>, <‘G’, 1>,

P5

<‘L’, 266>, <‘H’, 223>, <‘Shift+H’, 164>, <‘Tab’, 94>,
<‘Shift+L’, 83>, <‘F’, 35>, <‘E’, 19>, <‘Shift+Tab’, 8>,
<‘G’, 1>, <‘O’, 1>


    Path-Aware: For each GS, PAN analyzes the path from the

    beginning of the page to the beginning of that GS, and then injects a node that can be reached with the least number of repeated shortcut presses, using one of the Specific shortcuts in the user’s profile. If 𝑟 : 𝑠,𝑔 → N is a function that, assuming 𝐿(𝑠) is injected, indicates the number of times a shortcut 𝑠 ∈ 𝑆 has to be repeatedly pressed to navigate to the GS 𝑔 from the beginning of the page, then the annotation node 𝑦 injected is given by:

    Table 2: Shortcut profiles of participants. Only the Specific shortcuts have been shown as Basic shortcuts are not used for personalized annotation.

4.1 Evaluation Procedure and Results

The performance was measured in terms of the number of shortcut presses needed to reach the different GSs on a webpage.

4.1.1 Annotation vs. No Annotation. To measure the annotationinduced reductions in the number of shortcut presses for the GSs, we defined a metric 𝑑 (𝑔,𝑆) as the serial shortcut distance (i.e., number of Basic shortcuts) between the root of GS 𝑔 and the closest HTML element 𝑘 that is reachable with one of the Specific shortcuts 𝑠 in the user’s profile 𝑆.

Observe that 𝑑 indicates the average number of HTML elements ‘skipped’ (or equivalently shortcuts ‘saved’) due to annotation, which otherwise have to be traversed linearly one-by-one using Basic shortcuts, i.e., arrow keys. Also, while designing this metric we assumed an idealistic scenario where the user is aware of the optimal (least cost) navigation paths to all GSs with respect to the shortcuts in their profiles; therefore both path-agnostic and path-aware personalized annotation methods will yield the same results, which also represent the lower bound on performance gains achieved with personalized annotation.

4.1.2 Path-Agnostic vs. Path-Aware. Equation 1 presents the metric 𝑚 to compare the two annotation methods, where 𝑟 (𝑠,𝑔) is a func tion that, assuming 𝐿(𝑠) is injected for GS 𝑔, indicates the number of times a shortcut 𝑠 ∈ 𝑆 has to be repeatedly pressed to navigate to the GS 𝑔 from the beginning of the page, 𝑠𝑤 = 𝐿[-1] (𝑦 (𝑔)) for the Path-Aware annotation method, and 𝑠𝑔 = 𝐿[-1] (𝑦 (𝑔)) for the Path-Agnostic annotation method.

𝑚 = 𝑟 (𝑠𝑤,𝑔) − 𝑟 (𝑠𝑔,𝑔) (1)

The rationale behind selecting this metric was based on observations in related literature [5, 13] as well as observations during our own data collection phase described earlier, where we noticed that blind screen-reader users prefer to keep pressing the same shortcut multiple times, and switch shortcuts only if necessary. We also noticed that the users typically press Specific shortcuts first to skip large portions of the content, and switch to Basic shortcuts only if they are unable to find the desired segment using their preferred Specific shortcuts.

∀𝑔 ∈ 𝐺, 𝑦 (𝑔) = 𝐿(𝑥), where 𝑥 = argmin

𝑠 ∈𝑆

𝑟 (𝑠,𝑔)

While the Path-Agnostic annotation makes ‘local’ decisions based only on the HTML elements at the roots of GS subtrees, the PathAware annotation analyzes the ‘global’ HTML structure, with spe cific focus on the lengths of the serial paths from the beginning of the page to the roots of GS subtrees. The effectiveness of both these annotation techniques were evaluated as described next.

4 EVALUATION

As opposed to generating artificial shortcut profiles to evaluate our annotation techniques, we instead collected real shortcut profiles by recruiting 5 completely blind screen-reader users. These participants were recruited through local mailing lists and word-of-mouth. Initial screening was performed via phone interviews to enforce the inclusion criteria requiring familiarity with Web browsers and JAWS screen reader. All selected participants indicated that they were frequent users of web browsers, and that they spent at least 2-4 hours per day browsing the web. These participants were instructed to perform typical web-browsing tasks on a specific set of representative websites, and their shortcut presses were recorded with their consent. The received shortcut logs were then analyzed to build the corresponding shortcut profiles.

Table 2: Shortcut profiles of participants. Only the Specific shortcuts have been shown as Basic shortcuts are not used for personalized annotation.

Table 2 shows the details of the shortcut profiles collected from the participants. It can be clearly observed that no participant was aware of the special shortcuts ‘Q’ and ‘R’ corresponding to ARIA landmarks. Also, all participants knew the Basic Arrow shortcuts for navigating element-by-element. Interesting, only P4 knew about the Find shortcut (‘Ctrl+F’), although he sparingly used it. When probed regarding this in a phone interview, P4 stated that he finds Find shortcut only useful on familiar webpages; on unfamiliar webpages, he often cannot find the desired GS easily as the find keywords do not match with the text content in that GS.

(a)

(b)

Figure 2: Evaluation results (a) 𝑑[ˆ] - average reduction in shortcut presses with personalized annotation, and (b) 𝑚ˆ - average reduction in shortcut presses with Path-Aware annotation over Path-Agnostic annotation.

For each of the 5 collected profiles, we applied both the PathAgnostic and Path-Aware annotation methods on the test dataset

comprising 100 representative webpages (256 webpages, 948 total GSs) spanning several domains such as shopping, news, flight reservation, social media, classifieds, etc., and measured the average performances 𝑑[ˆ] and 𝑚ˆ over all 948 GSs from these websites.

Figure 2a presents the results for 𝑑[ˆ] indicating performance gains achieved by personalized annotation for each of the 5 participants. In reality, we expect annotations to yield significantly higher gains, as it is impractical for the users to remember the optimal paths for each GS on every webpage; they will likely take longer paths and press considerably more shortcuts without personalized annotation.

Figure 2b presents the results for 𝑚ˆ comparing the two annotation approaches. For all participants, on average, the performance of Path-Aware method was better than that of Path-Agnostic method. Also, for 256 out of 948 (∼ 27%) GSs in the test dataset, 𝑠𝑤 = 𝑠𝑔, i.e., the shortcuts from both the methods were the same thereby resulting in the same performance scores.

5 DISCUSSION AND FUTURE WORK

The results demonstrate the potential of personalized annotation in transparently improving the browsing experience of blind users, in terms of efficiency of screen-reader access and navigation. However, our preliminary evaluation also illuminated certain key areas of improvement that are necessary for our approach to be significantly more effective than the status quo. These are discussed next.

5.1 Limitations

This preliminary work only demonstrated the potential of personalized annotation in terms of expected reduction in shortcut presses, whereas measuring the exact reductions will require a within-subjects user study. In reality, it is also possible for PathAgnostic method to perform better than Path-Aware method as the

most frequently used shortcuts have higher chance of being repeatedly pressed first instead of the shortcuts chosen by the Path-Aware method. A user study is also required for comparison with other ‘non-personalized’ approaches mentioned earlier in the related lit erature. All these activities are scope of our future research.

5.2 Modification in addition to Injection

While adding new annotations to existing page DOM considerably improved the access effort for GSs on the page, we observed that the participants still had to navigate over multiple irrelevant nodes (albeit less than without annotations) before reaching the desired GS. For example, in one of the tasks, when P4 was trying to reach the main article content using his preferred ‘H’ shortcut, he had to press heading ‘H’ shortcut multiple times to first go over irrelevant advertisement blocks. This suggests that in addition to annotation, it is necessary to modify other parts of the webpage that are present inbetween these GSs. In the above example, changing the markup for the advertisements by replacing headings with some other elements (e.g., <p>) would have facilitated faster access to the article for P4.

5.3 Contextual Shortcut Preferences

In our approach, we considered the overall shortcut behavior of users in terms of their repertoire and frequency of usage. However, during evaluation, we observed that the preferences and frequencies varied according to the ‘type’ or ‘category’ of websites, e.g., shopping, booking flights, news, weather, searching, etc. This observation indicates that it is essential to incorporate the current context (type of website) as an additional influencing parameter while generating annotations for various GSs in webpages.

6 CONCLUSION

In this paper, we explored personalized semantics-driven webpage annotation techniques that strive to improve ease and efficiency of content access and navigation for blind screen-reader users. Contrary to the extant ‘one-size-fits-all’ techniques that require users to learn new interactive systems and also re-devise their browsing behavior, our approach offers a transparent solution that enables blind users to efficiently access/navigate content with their current personal screen-reading skills. The preliminary evaluation results demonstrate the potential and promise of our approach. As future work, we intend to explore in depth the multiple facets of blind users’ screen-reading behavior and incorporate them to transparently improve user experience and access efficiency for people who are blind.

Figure 2: Evaluation results (a) ˆ𝑑- average reduction in shortcut presses with personalized annotation, and (b) ˆ𝑚- average reduction in shortcut presses with Path-Aware annotation over Path-Agnostic annotation.

REFERENCES

[1] NV Access. 2020. NV Access. https://www.nvaccess.org/.

[2] Julian Alarte, David Insa, and Josep Silva. 2017. Webpage Menu Detection Based

on DOM. In SOFSEM 2017: Theory and Practice of Computer Science, Bernhard Steffen, Christel Baier, Mark van den Brand, Johann Eder, Mike Hinchey, and Tiziana Margaria (Eds.). Springer International Publishing, Cham, 411–422.

[3] Chieko Asakawa and Hironobu Takagi. 2000. Annotation-Based Transcoding for

Nonvisual Web Access. In Proceedings of the Fourth International ACM Conference on Assistive Technologies (Assets ’00). Association for Computing Machinery, New York, NY, USA, 172–179. https://doi.org/10.1145/354324.354588

[4] Chieko Asakawa, Hironobu Takagi, Shuichi Ino, and Tohru Ifukube. 2002. Au ditory and Tactile Interfaces for Representing the Visual Effects on the Web. In Proceedings of the Fifth International ACM Conference on Assistive Technologies (Edinburgh, Scotland) (Assets ’02). Association for Computing Machinery, New York, NY, USA, 65–72. https://doi.org/10.1145/638249.638263

[5] Vikas Ashok, Yury Puzis, Yevgen Borodin, and I.V. Ramakrishnan. 2017. Web

Screen Reading Automation Assistance Using Semantic Abstraction. In Proceedings of the 22nd International Conference on Intelligent User Interfaces (Limassol, Cyprus) (IUI ’17). Association for Computing Machinery, New York, NY, USA, 407–418. https://doi.org/10.1145/3025171.3025229

[6] Ali Selman Aydin, Shirin Feiz, Vikas Ashok, and IV Ramakrishnan. 2020. SaIL:

Saliency-Driven Injection of ARIA Landmarks. In Proceedings of the 25th In- ternational Conference on Intelligent User Interfaces (Cagliari, Italy) (IUI ’20). Association for Computing Machinery, New York, NY, USA, 111–115. https: //doi.org/10.1145/3377325.3377540

[7] Sean Bechhofer, Simon Harper, and Darren Lunn. 2006. SADIe: Semantic An notation for Accessibility. In Proceedings of the 5th International Conference on The Semantic Web (Athens, GA) (ISWC’06). Springer-Verlag, Berlin, Heidelberg, 101–115. https://doi.org/10.1007/119260788

[8] Jeffrey P. Bigham, Jeremy T. Brudvik, and Bernie Zhang. 2010. Accessibility by

Demonstration: Enabling End Users to Guide Developers to Web Accessibility Solutions. In Proceedings of the 12th International ACM SIGACCESS Conference on Computers and Accessibility (Orlando, Florida, USA) (ASSETS ’10). Association for Computing Machinery, New York, NY, USA, 35–42. https://doi.org/10.1145/ 1878803.1878812

[9] Jeffrey P. Bigham and Richard E. Ladner. 2007. Accessmonkey: A Collaborative

Scripting Framework for Web Users and Developers. In Proceedings of the 2007 International Cross-Disciplinary Conference on Web Accessibility (W4A) (Banff, Canada) (W4A ’07). Association for Computing Machinery, New York, NY, USA, 25–34. https://doi.org/10.1145/1243441.1243452

[10] Jeffrey P. Bigham, Tessa Lau, and Jeffrey Nichols. 2009. Trailblazer: Enabling Blind

Users to Blaze Trails through the Web. In Proceedings of the 14th International Conference on Intelligent User Interfaces (IUI ’09). Association for Computing Ma- chinery, New York, NY, USA, 177–186. https://doi.org/10.1145/1502650.1502677

[11] Syed Masum Billah, Vikas Ashok, Donald E. Porter, and I.V. Ramakrishnan. 2017.

Speed-Dial: A Surrogate Mouse for Non-Visual Web Browsing. In Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility. ACM, 3132531, 110–119. https://doi.org/10.1145/3132525.3132531

[12] Michael Bolin, Matthew Webber, Philip Rha, Tom Wilson, and Robert C. Miller.

    Automation and Customization of Rendered Web Pages. In Proceedings

of the 18th Annual ACM Symposium on User Interface Software and Technology (Seattle, WA, USA) (UIST ’05). Association for Computing Machinery, New York, NY, USA, 163–172. https://doi.org/10.1145/1095034.1095062

[13] Yevgen Borodin, Jeffrey P. Bigham, Glenn Dausch, and I. V. Ramakrishnan. 2010.

More Than Meets the Eye: A Survey of Screen-reader Browsing Strategies. In Proceedings of the 2010 International Cross Disciplinary Conference on Web Acces- sibility (W4A) (Raleigh, North Carolina) (W4A ’10). ACM, New York, NY, USA, Article 13, 10 pages. https://doi.org/10.1145/1805986.1806005

[14] Andy Brown and Simon Harper. 2013. Dynamic Injection of WAI-ARIA into Web

Content. In Proceedings of the 10th International Cross-Disciplinary Conference on Web Accessibility (Rio de Janeiro, Brazil) (W4A ’13). Association for Computing Machinery, New York, NY, USA, Article 14, 4 pages. https://doi.org/10.1145/ 2461121.2461141

[15] Deng Cai, Shipeng Yu, Ji-Rong Wen, and Wei-Ying Ma. 2004. VIPS: A vision based

page segmentation algorithm. Report. Microsoft technical report.

[16] Prathik Gadde and Davide Bolchini. 2014. From Screen Reading to Aural Glancing:

Towards Instant Access to Key Page Sections. In Proceedings of the 16th International ACM SIGACCESS Conference on Computers & Accessibility (Rochester, New York, USA) (ASSETS ’14). Association for Computing Machinery, New York, NY,

[17] Darren Guinness, Edward Cutrell, and Meredith Ringel Morris. 2018. Caption

Crawler: Enabling Reusable Alternative Text Descriptions Using Reverse Image Search. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–11. https://doi.org/10.1145/3173574.3174092

[18] Apple Inc. 2020. Vision Accessibility - Mac - Apple. https://www.apple.com/

[19] Jonathan Lazar, Aaron Allen, Jason Kleinman, and Chris Malarkey. 2007. What

Frustrates Screen Reader Users on the Web: A Study of 100 Blind Users. Interna- tional Journal of human-computer interaction 22, 3 (2007), 247–269.

[20] Gilly Leshed, Eben M. Haber, Tara Matthews, and Tessa Lau. 2008. CoScripter:

Automating & Sharing How-to Knowledge in the Enterprise. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Florence, Italy) (CHI ’08). Association for Computing Machinery, New York, NY, USA, 1719–1728. https://doi.org/10.1145/1357054.1357323

[21] Ian Li, Jeffrey Nichols, Tessa Lau, Clemens Drews, and Allen Cypher. 2010. Here’s

What i Did: Sharing and Reusing Web Activity with ActionShot. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Atlanta, Georgia, USA) (CHI ’10). Association for Computing Machinery, New York, NY, USA, 723–732. https://doi.org/10.1145/1753326.1753432

[22] Valentyn Melnyk, Vikas Ashok, Yury Puzis, Andrii Soviak, Yevgen Borodin,

and I. V. Ramakrishnan. 2014. Widget Classification with Applications to Web Accessibility. In Web Engineering, Sven Casteleyn, Gustavo Rossi, and Marco Winckler (Eds.). Springer International Publishing, Cham, 341–358.

[23] Paula Montoto, Alberto Pan, Juan Raposo, Fernando Bellas, and Javier López.

    Automating Navigation Sequences in AJAX Websites. In Web Engineering,

Martin Gaedke, Michael Grossniklaus, and Oscar Díaz (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 166–180.

[24] Christopher Power, André Freire, Helen Petrie, and David Swallow. 2012. Guide lines Are Only Half of the Story: Accessibility Problems Encountered by Blind Users on the Web. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Austin, Texas, USA) (CHI ’12). ACM, New York, NY, USA, 433–442. https://doi.org/10.1145/2207676.2207736

[25] Jyotika Prasad and Andreas Paepcke. 2008. Coreex: Content Extraction from

Online News Articles. In Proceedings of the 17th ACM Conference on Infor- mation and Knowledge Management (Napa Valley, California, USA) (CIKM ’08). Association for Computing Machinery, New York, NY, USA, 1391–1392.

[26] Yury Puzis, Yevgen Borodin, Rami Puzis, and I.V. Ramakrishnan. 2013. Predictive

Web Automation Assistant for People with Vision Impairments. In Proceed- ings of the 22nd International Conference on World Wide Web (Rio de Janeiro, Brazil) (WWW ’13). Association for Computing Machinery, New York, NY, USA, 1031–1040. https://doi.org/10.1145/2488388.2488478

[27] Yury Puzis, Yevgen Borodin, Andrii Soviak, Valentyn Melnyk, and I. V. Ra makrishnan. 2015. Affordable Web Accessibility: A Case for Cheaper ARIA. In Proceedings of the 12th Web for All Conference (Florence, Italy) (W4A ’15). Association for Computing Machinery, New York, NY, USA, Article 32, 4 pages. https://doi.org/10.1145/2745555.2746657

[28] Freedom Scientific. 2020. JAWS [®] – Freedom Scientific. http://www. freedomscientific.com/products/software/jaws/.

[29] Shaomei Wu, Jeffrey Wieland, Omid Farivar, and Julie Schiller. 2017. Automatic

Alt-Text: Computer-Generated Image Descriptions for Blind Users on a Social Network Service. In Proceedings of the 2017 ACM Conference on Computer Supported Cooperative Work and Social Computing (Portland, Oregon, USA) (CSCW ’17). Association for Computing Machinery, New York, NY, USA, 1180–1192.

[30] Yanhong Zhai and Bing Liu. 2005. Web Data Extraction Based on Partial Tree

Alignment. In Proceedings of the 14th International Conference on World Wide Web (Chiba, Japan) (WWW ’05). Association for Computing Machinery, New York, NY, USA, 76–85. https://doi.org/10.1145/1060745.1060761

[31] Jun Zhu, Zaiqing Nie, Ji-Rong Wen, Bo Zhang, and Wei-Ying Ma. 2006. Simultane ous Record Detection and Attribute Labeling in Web Data Extraction. In Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Philadelphia, PA, USA) (KDD ’06). Association for Computing Ma- chinery, New York, NY, USA, 494–503. https://doi.org/10.1145/1150402.1150457

[32] Manuel Álvarez, Alberto Pan, Juan Raposo, Fernando Bellas, and Fidel Cacheda.

    Finding and extracting data records from web pages. Journal of Signal

Processing Systems 59, 1 (2010), 123–137.

---

This full-text Seed edition preserves attribution to the ACM Hypertext 2020 version of record. Formatting was converted from that version under the supplied ACM authorization.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime