Bitget App
Trade smarter
Buy cryptoMarketsTradeFuturesEarnSquareMore
New initiative enhances AI access to Wikipedia information

New initiative enhances AI access to Wikipedia information

Bitget-RWA2025/10/01 13:25
By:Bitget-RWA

On Wednesday, Wikimedia Deutschland revealed a new database designed to make Wikipedia’s extensive information more easily available to AI systems.

Named the Wikidata Embedding Project, this platform utilizes a vector-based semantic search method—a process that enables computers to interpret the meanings and connections between words—on the vast data from Wikipedia and its related sites, which together hold close to 120 million records.

By integrating support for the Model Context Protocol (MCP)—a standard that enables AI to interact with data sources—the initiative allows LLMs to access the data through natural language queries more effectively.

Wikimedia’s German division developed the project in partnership with neural search company Jina.AI and DataStax, a real-time data training firm owned by IBM.

For years, Wikidata has provided machine-readable information from Wikimedia sites, but previous tools only supported keyword searches and SPARQL, a specialized query language. The updated system is better suited for retrieval-augmented generation (RAG) setups, which let AI models incorporate external knowledge, giving developers the ability to anchor their models in content reviewed by Wikipedia editors.

The data is organized to deliver essential semantic context. For example, searching for “scientist” in the database will yield lists of notable nuclear scientists, scientists affiliated with Bell Labs, translations of “scientist” in various languages, an approved Wikimedia image of scientists at work, and related terms like “researcher” and “scholar.”

Anyone can access the database on Toolforge. Additionally, Wikidata will host a webinar for developers interested in the project on October 9th.

This initiative arrives at a time when AI developers are urgently seeking reliable, high-quality data to refine their models. Training environments have grown more advanced—often built as intricate systems rather than simple datasets—but they still depend on carefully curated information. For applications demanding high precision, trustworthy data is crucial. While Wikipedia may have its critics, its content is far more fact-based than broad collections like Common Crawl, which aggregates vast numbers of web pages from the internet.

Sometimes, the pursuit of top-tier data can be costly for AI companies. For instance, in August, Anthropic agreed to pay $1.5 billion to settle a lawsuit with a group of authors whose works were used for training, resolving all related claims.

In a statement to the media, Wikidata AI project manager Philippe Saadé highlighted the project’s independence from major tech firms or leading AI labs. “The launch of this Embedding Project demonstrates that advanced AI doesn’t need to be dominated by a few corporations,” Saadé said. “It can be open, collaborative, and designed to benefit everyone.”

0

Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.

PoolX: Earn new token airdrops
Lock your assets and earn 10%+ APR
Lock now!

You may also like

Bitcoin News Today: Bitcoin's Major Holders Selling Challenges ETF Support at $90k

- Bitcoin whale inflows hit 9,000 BTC on Nov 21, 2025, with 45% of deposits from large holders, signaling intensified selling pressure amid a seven-month price drop to $80,600. - Exchange inflows surged to $40B weekly, with Binance’s stablecoin reserves reaching $51B, reflecting capital shifts toward dollar-pegged assets amid market uncertainty. - ETF inflows (e.g., BlackRock’s IBIT) provided limited counterbalance, totaling $21M on Nov 27, contrasting with earlier $903M outflows and whale-driven altcoin d

Bitget-RWA2025/11/30 07:58
Bitcoin News Today: Bitcoin's Major Holders Selling Challenges ETF Support at $90k

Solana News Today: Crypto at a Turning Point—Speculation Mania or Institutional Domination?

- Arthur Hayes, ex-BitMEX CEO, boosted DeFi exposure with 2.01M ENA and 33K ETHFI tokens amid crypto volatility. - Solana (SOL) struggles to break $150, forming a bear flag pattern that could trigger a 30% drop to $99 if $140 support fails. - Nasdaq's IBIT options proposal and Grayscale's Zcash ETF filing signal growing institutional crypto adoption amid fragmented market dynamics. - Astra Bitcoin's hybrid model blends TradFi/DeFi assets to address volatility concerns, yet speculative momentum remains evid

Bitget-RWA2025/11/30 07:40
Solana News Today: Crypto at a Turning Point—Speculation Mania or Institutional Domination?

Bitcoin Updates: With Retail Investors Declining, Large Holders and ETFs Influence Bitcoin's Direction

- Bitcoin's $91,000 rebound highlights institutional dominance over retail traders, driven by ETF inflows and whale accumulation. - Bhutan's $970,000 ETH staking and RGB20 protocol advancements signal institutional validation of Bitcoin's programmable finance potential. - Solana's $8.2M ETF outflow and $36M hack contrast Bitcoin's stability, as large holders buffer against volatility. - ETF-driven price dynamics and privacy-focused products like Zcash ETFs reflect shifting market structure toward instituti

Bitget-RWA2025/11/30 07:40
Bitcoin Updates: With Retail Investors Declining, Large Holders and ETFs Influence Bitcoin's Direction

Zcash Latest Updates: Zcash ETF Anticipation Faces Bearish Trends—Will This Privacy Coin Overcome the Downturn?

- Zcash (ZEC) nears critical $442.53 support as technical indicators signal bearish momentum with 12/12 "Strong Sell" signals. - Grayscale's proposed ZCSH ETF aims to institutionalize privacy-focused crypto access, holding 394,400 ZEC valued at $199M. - Market remains muted despite ETF filing, with ZEC down 1.4% amid regulatory uncertainty and broader crypto volatility. - ETF approval could boost ZEC liquidity like Bitcoin ETFs, but traders watch $442.53 support and SEC review outcomes.

Bitget-RWA2025/11/30 07:40