Tagging NCSES Website Content

Image
Man at desktop computer

The National Center for Science and Engineering Statistics (NCSES), part of the National Science Foundation, serves as the principal source of analytical and statistical reports, data, and related publications that provide insight into the nation’s science and engineering resources. 

To enhance user experience and ensure efficient information retrieval, NCSES sought to maximize the discoverability of its web content through a comprehensive tagging initiative. The project aimed to develop a topic-driven categorization and organization system for NCSES content, enhancing the ability of users to find and discover information on the newly redesigned NCSES website.

The Challenge

The NCSES website suffered from poor user experience due to inadequate content tagging. Because of the lack of complete and accurate tagging, users struggled to locate the information they were seeking. The lack of a good tagging system also hindered the website’s ability to suggest related content based on user interests. NCSES recognized the need to improve its website and saw an opportunity during a strategic website redesign and technology upgrade to implement a new, robust tagging system. However, it faced a challenge due to thousands of pieces of website content that were not adequately tagged, and it had no system or interface to support content tagging.

Our Role

AIR, along with our subcontractor Vertivis, was engaged by NCSES to use applied data science and content management methods to tag thousands of pieces of web content by topic and areas of interest in order to enable users to find information on the NCSES website.

Our initial assessment phase involved extensive discussions with the NCSES technology team and subject matter experts (SMEs). These interactions helped us understand the best methods for extracting content from the existing website, including archived legacy content. We developed a business process that facilitated broader SME input into the content-tagging process.

A significant realization during implementation was the necessity of creating a Tagging Review Portal. This web-based interface allowed non-technical SMEs to review, revise, and augment the automated tags generated by our data science methods. Managing the development of this portal with our subcontractor was a critical part of our role.

 

Methodology

Our methodology was divided into two main phases: Content Gathering and Tagging Methods.

Content Gathering: We utilized web-scraping techniques to extract text and metadata for legacy NCSES web products. This process involved developing customized scrapers for different content types and ensuring comprehensive data extraction. The extracted data were then compared with the master content file provided by NCSES to ensure completeness.

Tagging Methods: We employed a combination of rule-based and model-based tagging approaches. For tables and figures, we used rule-based methods, developing keywords and tagging rules to classify content accurately. For text-based content, we used machine learning and natural language processing (NLP) models to achieve precise tagging through topic classification. The tagging process was iterative, involving extensive testing and refinement with SME feedback.

 

Outcome

The success of the project was measured using key metrics such as accuracy, precision, and recall for the tagging models. Our models achieved excellent performance metrics, ensuring high-quality tagging. The Tagging Review Portal facilitated SME review, allowing for real-time feedback and iterative improvement of the tagging process.

The current NCSES website now uses the topic and area of interest tags we developed, significantly enhancing search functionality and user experience. Additionally, we documented lessons learned to inform future efforts and provided NCSES with the capability to replicate the tagging process for new content.

Overall, the project not only addressed the immediate challenges faced by NCSES but also laid the groundwork for ongoing content management and discoverability improvements, enhancing the overall user experience on the NCSES website. Key highlights of the project include:

  1. Effective Application of Advanced Methods: We applied well-established large language model (LLM) and NLP techniques to develop an automated tagging model. This ensured accurate and efficient categorization of a wide array of NCSES content.
  2. Emphasis on Quality Assurance: Throughout the project, we maintained a strong focus on quality assurance. This included rigorous code reviews, comprehensive content analysis, and iterative feedback from SMEs to continually refine and improve the tagging accuracy.
  3. Effective Web Scraping: Our team successfully executed web scraping techniques to extract all necessary content from the existing NCSES website, including complex and archived legacy content. This thorough extraction process was critical for the subsequent tagging efforts.
  4. User-Friendly Review Process: We developed an intuitive web-based interface that enabled SMEs to easily review, revise, and augment the automated tags. This collaborative approach ensured the final tagging was accurate and met the needs of NCSES.
  5. Seamless Collaboration: We worked closely with the NCSES technical team to ensure all tags were seamlessly integrated into their new content management system. This collaboration was essential for the successful implementation and functionality of the new tagging system.

     

Contact
Christina Jones

Christina Jones

Principal Data Scientist