No Longer Scraping By
How Congress Can Enable Better Constituent Information-Gathering in the AI Age
The Modernization and Innovation Subcommittee of the Committee on House Administration held a hearing Wednesday to understand the current landscape of legislative branch data availability and use. As Chair Stephanie Bice noted during the hearing, information about congressional activity, especially related to legislation, is more available than ever before. Artificial intelligence is adding an additional layer of complexity, however, to longstanding challenges in accessing legislative information, further undermining institutional trust. (We previewed the written testimony in a recent First Branch Forecast.)
SOME CONTEXT
The growth in the amount of information available is because of the significant progress legislative branch offices have made in adopting common structured data standards that now feed different systems and platforms. Although the Library of Congress launched Congress.gov in 2012, uniting internally-accessible data with a public-facing website to replace the old THOMAS.gov, users still could not download data that was technically public information on their own. Following public pressure, House appropriators authorized the creation of the Bulk Data Task Force that same year, bringing together the House Clerk, Congressional Research Service (which manages Congress.gov), the Government Publishing Office and other relevant offices to coordinate on making their data available via bulk download. Just as importantly, the task force began holding open quarterly meetings with members of the public so they could engage with legislative branch stakeholders.
The first major milestone for this public-private feedback mechanism came in 2014, when Senate offices relented and agreed to join the bulk data effort. With cross-chamber cooperation established, and public engagement continuing with relevant committees, Congress finally began to provide information about the status of bills as bulk data the following year. More data progressively became available through the years, and in 2022 the task force was renamed the Congressional Data Task Force to mark the progress.
Ultimately, public stakeholders succeeded in pushing the Library to provide a public API for Congress.gov data that launched in the fall of 2022. Demand certainly was there for it, as the Library shared that the API had received 1.3 billion requests at the Congressional Hackathon last September. It also has progressively added more information resources to the website, including some Congressional Research Reports, committee videos, and other content requested through public engagement with the task force.
LEGISLATIVE BRANCH DATA IN THE AGE OF AI
Artificial intelligence platforms are scrambling the public availability situation again. AI agents increasingly are scraping Congress.gov because it is the best reliable source of legislative branch information on the internet. It’s becoming too much of a good thing for the Library’s IT infrastructure. At the hearing, Library of Congress Deputy Chief Information Officer John Rutledge related that on June 16, AI bot traffic increased to the point the website started to crash and the Library had to institute human verification to maintain service. When it did so, 98% of the traffic dropped off. But non-human traffic is coming from people trying to build AI-powered tools to make the underlying data more accessible to people trying to understand congressional activity, so it’s not malicious.
The problem Rutledge relayed to the subcommittee is complicated because although there’s a reasonable sense that human eyeballs should get priority over scrapers, Congress should want AI models to use authoritative legislative branch data. Part of it is that the Library has not completed migrating Congress.gov to the cloud, because of tight resources. (That move is all the more essential given the exploding costs of servers, as Rutledge noted.) But most legislative branch entities are still thinking in terms of websites, which the House Clerk Office’s Kirsten Gullickson (who is also the Congressional Data Task Force coordinator) noted in her oral testimony needs to evolve to acknowledge “we are publishing information into a digital ecosystem” being shaped by emergent tech. Charlotte Lee of the Customer Experience Leadership Institute emphasized the importance of well-structured data and well-documented APIs as keys to keeping AI from gumming up the human user experience.
Some parts of the legislative branch are building out that digital ecosystem to meet changing technology. Late last year, the Government Publishing Office announced it was launching a Model Context Protocol server to help AI platforms keep up to date on federal government activity via the authoritative repository of documents.
Although it is creating challenges for Congress.gov at the moment, AI also holds some potential to solve a massive user experience problem for the Library and by extension the legislative branch. Congress.gov has so much information that it’s unwieldy even for those well versed in congressional operations and procedure. Searches return giant lists of hyperlinks. Typing the URL into a web browser isn’t even the way that the vast majority of users arrive at Congress.gov: Rutledge said that 95% of the site’s traffic is directed from an outside source as people are clicking on links surfaced by a search engine to find bills. More than 70% of congressional users of the site also come via search.
As Ranking Member Norma Torres noted, legislative data’s prioritization of accuracy, using the exact title or bill number in listing a piece of legislation, is not the way most people who come across something they want to look it up. They’re searching based on its content or issue area, which is exceedingly difficult without an interpretive layer to help identify what a user really is looking for and pointing them to the right source. Gullickson suggested additional use of semantic search as a solution. This challenge is heightened for legislative tracking, however, when a piece of legislation is included in a different bill. In the current version of Congress.gov, it essentially disappears from the user.
Bill tracking is something folks have been working on for awhile, though. Our Daniel Schuman built an open-source prototype called BillMap some years ago. Within the Library, CRS has developed a new version of a bill text analysis tool of its own, which it demonstrated at the last CDTF meeting. The Library doesn’t think it is properly “authoritative” enough, however, to release to the public yet. That’s ultimately a user experience decision, and Lee (who was part of the BillMap project) suggested it could be approached differently. She recalled working on the Clerk’s comparative print suite project, which prompted human users to verify changes in law that the system could not compute. The Library could set the parameters for users to make their own decisions with the information returned. These design decisions are necessary and something offices should be comfortable making because, as Lee noted, 100% accuracy with AI systems is impossible.
The Library has begun a major overhaul of Congress.gov, starting with two user interview sessions with 15 participants each. Rutledge said the goal is to have the redesign completed next year, with the ability to provide “personalized experience” based on different users’ needs. The next comment opportunity for this process will be on September 24 when the Library holds its annual Congress.gov public forum.
BEYOND CONGRESS.GOV
Of course, Congress.gov is only one of many congressional websites with public-facing resources – Torres noted there were 625 with a house.gov domain alone. Gullickson called for more online “wayfinding” to guide users to the appropriate digital resource. This principle is particularly important because the legislative branch is complex. She recommended that house.gov provide that role in helping people find authoritative information without first needing to understand the structure of the chamber.
Lee told the subcommittee that the legislative branch should develop its own AI layers not only to provide data analysis, but to be a responsive guide to the myriad of user types coming to its online resources. “The digital infrastructure should guide each along their path with clarity and context at every step to speak, be the primary, trusted source of truth wherever dialogue is happening,” she said. That includes “neutral explanation layers” built on the legislative branch’s authoritative sources that can provide better answers than AI models pulling from opinionated internet content. It’s a question of equal access to public information as high-priced subscription services provide information synthesis and analysis that the general public cannot access.
Ultimately, Lee concluded, Congress needed to be thinking about using legislative data to give itself a “voice” in AI-enabled public discovery and discussion of its activity because without it, the institution is silent and others fill in the blanks. Currently, it puts that activity – bill text, amendments, spending, votes – up on websites and hopes people visit and find them. This is all the more problematic when more people are only reading the AI summary at the head of their Google search and not clicking through, leaving Gemini to stand for Congress’s voice. Lee mentioned the United Kingdom’s Parliament even has a communications office.
These transformations are readily attainable and can be done with a lot of the talent the House already has in its offices. Even without them, Bice noted, staff are trying their own projects to create resources out of disparate data sources. But a major theme of the hearing was it will take additional resources to spin up innovation cycles. Gullickson spoke of a lack of dedicated staff capacity to collect the good ideas from legislative branch staff and route them to the right development team. (AGI submitted testimony to Legislative Branch Appropriations requesting CDTF be given funding to hire exactly this type of coordinator). The Library of Congress needs additional funds to complete the cloud transfer for Congress.gov and develop its own in-house AI model to assist with data analysis projects in a closed environment. The user experience research Lee championed takes resources and people. The hearing was a strong step in making the case for these types of investments, which the current version of the FY 2027 Legislative Branch Appropriations bill mostly leaves out.
GOING META
I would have liked to have attended the hearing but couldn’t because of a scheduling conflict. Experienced government information researcher Charlie Amiot also tried watching the hearing online as she lives on the West Coast. On her substack, she turned the experience into a meta-narrative about legislative data access for a hearing on this very topic, which is very illuminating and also a very different experience than mine. How two people looking for the same thing can go down two different paths because of different decisions on where to look felt like we were playing Zork.
To keep tabs on upcoming hearings, I regularly visit docs.house.gov and look at its weekly calendar. I also go to the committee website to see if they’ve posted a hearing notice – sometimes after what appears from the Clerk. I know that committee websites are maintained by the majority party in the House or Senate. The minority has a totally different website, the link to which is usually buried at the bottom. It’s up to majority staff to post committee records, documents, and alerts.
I want to watch the livestream of the hearing for a bit before I have to stop, so I go to the (majority) committee page, click on the title of the hearing, and load this page.
The link to the livestream is on the word “here.” Do you see it? For a time Wednesday, it either wasn’t posted yet or I forgot where to look and I got a little frustrated. But it took me directly to YouTube and the feed of the hearing. It also will serve as an unedited record of the hearing along with the video posted on Congress.gov with the caveat, “availability of video proceedings varies by committee.”
I wanted to have a transcript of the hearing to write this summary. The House doesn’t have a rapid transcription service available yet. I know there will be an official transcription at some point on Congress.gov, but it takes awhile to produce. I have to put out a newsletter for Monday, so I copy the link into a transcriber service AGI pays for. Now, I’m ready to write.
Amiot’s user experience starts with her decision to find the hearing video via CHA’s YouTube channel. I didn’t even know that existed to be honest with you. Thus begins her user odyssey with just trying to find what part of YouTube has the livestream. The committee channel is cluttered with clips and other hearings and livestreams don’t appear under the “video” tab.
She, too, wanted a full transcript, but was blocked by her tool because it thought the content was restricted. So she had to transcribe nine different clips from the YouTube channel. Those clips, however, were produced by the majority committee staff, who again are in charge of committee website content. Because the committee website is the voice of the majority party, all of the clips are questions from Chair Bice. (Some other channel clips are members’ interviews on cable TV.) Amiot has to watch the full hearing video to hear from Ranking Member Torres.
Amiot posted this experience on her substack that alerts people to new CRS report releases (and she’s on the Depository Library Council) so she is hardly a novice with legislative information. Nevertheless, we had totally different user experiences for the same information based on different mental maps about congressional committees. Unfortunately, hers was a lot more time consuming. And be sure to read her summary for things I may have left out.
SORTING OUT MANDATED REPORTS
After a decade-long effort, Congress enacted the Access to Congressionally Mandated Reports Act requiring agencies to provide a duplicate copy of many of the reports they submit to Congress to the Government Publishing Office, which maintains them in an open repository. Although it is a victory for government transparency, the recordkeeping system behind mandated reports makes it challenging to know which reports are missing and how they link up to specific congressional requests.
To help sort things out until Congress makes fixes to ACMRA, the American Governance Institute has launched a technology project with government data and tech expert Dave Zvenyach to experiment with methods for sorting out the report collection and identifying what’s missing. Zvenyach presented preliminary findings of the project at the Congressional Data Task Force meeting in June, and he and Daniel recorded a more in-depth conversation about it recently that is posted on the Congressional Data Coalition blog.
Among Dave’s findings:
It is technologically possible to create a crosswalk between the reports identified in the Clerk’s report and those received by GPO.
Of the nearly 3,300 mandated reports, GPO has received less than 1/3, or approximately 1,060 filings.
GPO has received reports from agencies that are not in the Clerk’s list, approximately 400 reports so far.
MEMBER TRADING
The saga of forbidding members of Congress from trading stocks almost certainly is over for this session. After multiple bipartisan bills from the rank-and-file and a discharge petition pushed a ban into potential legislative reality, House leadership in both parties have stepped in to steer the issue back to their preferred position of allowing trading to continue. Democratic leadership acted first, adding bans on the President and Vice President to their members’ effort to spike Republican support. Last week, Republican leadership added their voter ID bill that has been DOA in the Senate to the Stop Insider Trading Act. It passed, garnering 13 Democratic votes.
Rep. Thomas Massie broke it down as only he can.
Senator Cynthia Lummis has reintroduced a bill that would ban federal officials, including members of Congress, from issuing or sponsoring digital assets. But the ban would start on inauguration day for the next administration.
TENURE
A group of younger, more conservative Democrats have proposed adopting term limits for committee leaders in the next Congress, citing the age of many current ranking members and the eight Democratic members who have died since 2023. Majority Democrats, which has 18 congressional members, released a short video proposing the rules change that featured recently deceased Agriculture ranking member David Scott prominently for struggling in his role. It also noted the political blowback younger members receive in challenging elder leaders, even if they are in physical and mental decline.
Generational turnover certainly is not happening in many committees, although the Democrat with the longest tenure serving as a chair or ranking member, Nydia Velazquez, is retiring after this term. She has been at the head of the House Small Business Committee for an astonishing 27.5 years. Five other House Democrats have been the party’s lead on committees for at least 11 years. Their average age is 75.8 years.
This long-term tenure is attributable to several factors. The Congressional Black Caucus and Congressional Hispanic Caucus historically have opposed caucus term limits because their members faced advancement barriers earlier in their careers. From a realist perspective, committee leadership positions are a form of power for these caucuses, and power is rarely given up voluntarily in Congress. But Rep. Nancy Mace’s ridiculous new resolution abolishing all affinity group caucuses (except, of course, the Anti-Woke Caucus) is indicative of the continuing concern that remains for non-white members that they will be shut out of leadership opportunities.
The Republican Conference rules do limit members to three consecutive terms as chair or ranking member. Arguably, these limits have accelerated the retirement of more senior members once reached. They do, however, provide regular turnover of gavels, which is a point of contrast with House Democrats, whose long-tenured chairs reflect a deeper tradition of allowing senior members to control committees for decades.
Daniel did some preliminary research on Democratic committee tenure over the last 40 years with the assistance of ChatGPT. It pulled from legislative branch data sources we would have used to create this graphic, but we have not had time to fact check it. It’s notable that the Small Business Committee has had only two Democratic leaders since 1987, while Energy and Commerce has had three. Rep. John Conyers held the Judiciary leadership position from 1995 until he finally resigned from Congress in 2017 after sexual harassment and hush money allegations.
We’ll dig into this topic more with additional research. Although we do not think term limits are the best reform to do so, some change is needed for the Democratic Caucus to change committees from personal fiefdoms and hold chairs and ranking members more accountable for their performance in leadership. It also would be an improvement institutionally if members with relevant subject-matter expertise or policy positions that reflected a majority of the caucus had better shots at ascending to committee leadership. Of course, that’s a question about how much control caucus leadership is permitted to hold and thus part of a more involved reform conversation not only for Democrats, but the entire House.




I’m East Coast, I just live like I’m in the Pacific timezone!
Dead on otherwise.
Thanks for this terrific summary of the hearing!