Historical Solana Data: What to Look For in a Solana Data Provider
Raw transactions, decoded protocol tables, RPC, APIs, and Parquet datasets solve different problems. Learn what historical Solana data contains and how to choose the right provider for analysis, backtesting, and machine learning.
If you need historical Solana data for backtesting, wallet analysis, market research, token analytics, or machine-learning models, the first question is usually:
Where can I get the Solana blockchain data?
The more important question is:
What kind of Solana data do you actually need?
There is a major difference between downloading raw transactions from an RPC node and receiving a structured dataset that already tells you that a transaction contained a Pump.fun buy, a Jupiter swap, a Raydium liquidity event, or a token transfer.
Both originate from the same blockchain. They are not equally useful for analysis.
This guide explains what historical Solana data contains, how Solana transactions become analytics datasets, and what to check before choosing a historical Solana data provider.
What is historical Solana data?
Historical Solana data is the record of activity that has already occurred on the Solana blockchain.
At the lowest level, this includes:
- blocks and slots
- transactions and signatures
- accounts and account state
- program instructions and inner instructions
- logs
- token balances and SOL balances
- timestamps
A Solana transaction can contain one or more instructions. Each instruction invokes a Solana program and specifies the accounts involved plus instruction-specific data. Transactions can also cause programs to invoke other programs through cross-program invocations.
That structure is powerful for applications but creates additional work for analysts.
A raw transaction might tell you which program was invoked and provide its instruction bytes. It does not necessarily hand you a convenient row saying:
Wallet X bought 14,250 tokens on Pump.fun for 1.72 SOL.
To produce that row, someone has to understand the relevant program, identify the instruction, decode its arguments and accounts, and turn the result into a stable schema.
That difference is at the center of choosing a historical Solana data provider.
Raw Solana blockchain data vs. decoded Solana data
There are roughly two levels of historical data you can work with.
Raw Solana data
Raw data stays close to what Solana itself exposes. For a transaction, you may have fields such as:
- signature
- slot and block time
- account keys
- program IDs
- compiled instructions
- instruction data
- transaction metadata
- pre- and post-token balances
- logs
This is useful when you need maximum control or when you are studying something for which no decoder exists. But raw data frequently moves the difficult work downstream.
If you want to answer “How many Pump.fun tokens were created last month?”, you first need to know which Pump.fun program and instruction represent a creation, account for instruction variants, decode them correctly, and decide what constitutes a successful creation.
If you want “What was the average PumpSwap trade size during the first hour after a token migrated?”, you need Pump.fun creation data, migration information, PumpSwap trading data, timestamps, token identities, and a way to join those records reliably.
Raw history gives you the evidence. It does not automatically give you the analytical table.
Decoded Solana protocol data
Decoded datasets move one level higher. Instead of only storing program_id, accounts, and instruction_data, the provider interprets the protocol and exposes fields relevant to that instruction or event.
For example, a decoded Pump.fun buy table can expose the token mint, bonding curve, user, amount, and maximum SOL cost rather than requiring the analyst to decode those values for every transaction.
datastore.sh Solana datasets follow this second approach: protocol instructions and events are delivered as typed tables in Parquet, with schemas documented before the data is purchased.
Why Solana data is not one universal table
A common mistake is thinking of “Solana blockchain data” as one giant transactions table.
For some questions, that works. For serious protocol analysis, it usually does not.
Consider just a few types of activity on Solana:
- a Pump.fun token creation
- a Pump.fun bonding-curve buy
- a PumpSwap swap
- a Jupiter routed trade
- a Raydium CLMM liquidity position
- a Meteora DLMM position change
- a Drift perpetual trade
- an SPL Token mint
- a Token-2022 transfer
They all happen on Solana, but they represent different programs with different instruction formats, account layouts, and semantics.
The same generic instruction_data field therefore means very different things depending on which program is processing it.
A useful historical Solana dataset preserves blockchain-level identifiers while also representing protocol-specific meaning.
What should Pump.fun historical data contain?
Suppose you are researching token launches. A useful Pump.fun dataset should let you distinguish operations such as:
- token creation
- buys and sells
- creator-fee operations
- relevant configuration changes
- associated protocol events
You should then be able to ask:
- How many tokens were created per day?
- Which wallets repeatedly bought tokens shortly after creation?
- How much SOL entered a bonding curve before migration?
- What percentage of newly created tokens reached a particular activity threshold?
- How did early buyers behave after migration?
- Which creators repeatedly launched tokens?
These are difficult questions to answer from a generic transaction table alone. A protocol-specific dataset exposes decoded instructions and events rather than only raw transaction payloads.
Why PumpSwap data is a separate problem
PumpSwap data becomes important once research moves beyond the initial Pump.fun bonding-curve stage.
Questions might include:
- What happened to liquidity after a token moved to PumpSwap?
- What was its first traded price?
- How much volume occurred during its first 10 minutes?
- Which wallets bought immediately after pool creation?
- How did reserves change?
- What were LP deposits and withdrawals?
- How much slippage would a historical order have experienced?
This requires information about PumpSwap’s own instructions and events. Treating Pump.fun and PumpSwap as exactly the same dataset can hide an important boundary in the token’s lifecycle.
A useful research pipeline can join the two through common identifiers such as token mints, pools, wallets, and timestamps while preserving which protocol generated each event.
What to look for in a Solana data provider
The phrase Solana data provider can describe very different products. Some providers primarily expose RPC infrastructure. Others provide APIs, indexed query platforms, warehouse feeds, or bulk blockchain datasets.
Start with the research question and then evaluate the delivery model.
1. Is the data raw or decoded?
Ask whether you receive something like:
program_id
accounts
instruction_data
logsor something closer to:
timestamp
signature
mint
user
amount
max_sol_cost
bonding_curveNeither is inherently better. If you are building your own universal Solana indexer, raw data may be appropriate. If you want to analyze Pump.fun trading behavior, decoded data removes a large amount of work.
2. What exactly does historical coverage mean?
“Historical Solana data” is not a useful coverage description by itself. Check:
- earliest available timestamp or slot
- latest available timestamp or slot
- whether coverage varies by protocol
- whether failed transactions are included
- whether inner instructions are captured
- whether historical program versions are supported
- whether data was backfilled or captured live
- whether gaps are documented
For protocol research, coverage needs to be defined at the dataset or table level rather than as a vague claim about the entire blockchain.
3. Is the schema documented before purchase?
Do not wait until terabytes of data arrive to discover what a column means. You should be able to inspect:
- table names
- column names and data types
- field descriptions
- sample records
- schema version
- partitioning
This is particularly important for protocol data because a field called amount can mean very different things across instructions. Browse the Solana catalog to inspect available datasets and schemas.
4. Can you keep the dataset?
There are two fundamentally different ways to consume historical blockchain data.
Query an external service
analysis → API request → provider → responseThis can be convenient for small or interactive requests.
Own the historical data
Parquet → DuckDB / Polars / ClickHouse / Spark / warehouse → analysisBulk files become particularly useful when you repeatedly scan the same historical range. If you are training 200 variations of a trading model against the same year of historical swaps, downloading the underlying dataset once is fundamentally different from requesting the same history from an API for every experiment.
The core historical product at datastore.sh is the second model: customers receive partitioned Parquet files that can be retained and queried in their own infrastructure rather than depending on a metered historical query API.
5. Can you verify what you received?
Historical datasets often become inputs to research that must be reproduced months later. Useful delivery metadata includes:
- dataset version
- schema version
- coverage period
- partition information
- file manifest
- SHA-256 checksums
Checksums let you confirm that files have not changed or become corrupted. Versioning also matters when a provider discovers a decoding problem later. Quietly changing an existing dataset makes previously generated results difficult to reproduce.
Historical Solana data for backtesting
Backtesting is one of the strongest use cases for bulk blockchain datasets.
Imagine you want to test this strategy:
- identify every Pump.fun token creation;
- measure activity during its first 30 seconds;
- enter when a set of conditions is reached;
- follow the token through its subsequent lifecycle;
- simulate exits using later trading data.
The difficult part is not writing if signal: buy(). The difficult part is building the historical state that existed when the decision would have been made.
That may require combining creation events, buys, sells, token balances, bonding-curve state, wallet activity, migration events, PumpSwap activity, and timestamps.
You also need to avoid future leakage. If your simulated strategy makes a decision at 12:00:05, every feature used by that decision must have been observable by 12:00:05.
A good historical blockchain dataset gives you the raw ingredients. A correct backtest still depends on how those ingredients are used.
Historical Solana data for wallet analysis
Wallet analysis requires a different view of the same blockchain history. Typical questions include:
- Which tokens did a wallet interact with?
- Which protocols does it use?
- How quickly does it buy after launch?
- What other wallets participate in the same tokens?
- Does it repeatedly appear before large price movements?
- How much volume has it generated?
- How long does it hold positions?
Protocol-decoded data makes these questions easier because the analyst can work with actual operations rather than repeatedly decoding generic transaction instructions.
Historical Solana data for machine learning
Machine-learning pipelines benefit from another property of bulk datasets: repeatability.
A model might extract features including trade counts, buy/sell imbalance, unique traders, wallet concentration, liquidity changes, token age, transaction frequency, creator activity, and historical wallet behavior.
Feature generation can require scanning the same history many times. Keeping the underlying dataset means the same immutable input can be reused while feature logic and models evolve. It also makes it easier to associate a training run with a specific dataset and schema version.
Why Parquet is useful for blockchain datasets
Blockchain history is naturally large. Storing analytical data in a columnar format such as Parquet has an important advantage: an analytical query often needs only a small fraction of the available columns.
Suppose a table contains 40 fields but your query requires only:
timestamp
mint
user
amountA columnar query engine can avoid reading many unrelated columns.
Parquet also integrates with DuckDB, Polars, ClickHouse, Spark, Snowflake, and BigQuery. That makes the dataset portable rather than tying analysis to one provider’s custom query language.
RPC, API, or blockchain dataset: which one do you need?
| Requirement | Usually best starting point |
|---|---|
| Current account state | RPC |
| Submit Solana transactions | RPC |
| Retrieve a small amount of history | RPC or API |
| Build an application around indexed data | API |
| Analyze millions or billions of historical records | Bulk dataset or warehouse |
| Run repeated historical backtests | Bulk dataset |
| Train models against protocol activity | Bulk dataset |
| Build your own decoder or indexer | Raw transaction or archive data |
| Analyze a known protocol | Decoded protocol dataset |
These categories overlap. The important part is matching the data delivery method to the workload rather than simply searching for the “best Solana API.”
Questions to ask before buying a historical Solana dataset
Before committing to a provider, ask:
- What exact period does the dataset cover?
- What tables are included?
- Are the records raw or decoded?
- Can I inspect the schema before purchasing?
- Can I see sample rows?
- How are protocol upgrades handled?
- Are failed transactions represented?
- How are corrections published?
- Is the data versioned?
- Can I verify files using checksums?
- Can I keep the dataset permanently?
- What identifiers are preserved for joining back to Solana?
If a provider cannot answer these questions clearly, it will be difficult to know exactly what a historical analysis represents.
Where can you get historical Solana data?
There are several ways to obtain Solana history. You can operate archival infrastructure and build your own indexing and decoding pipeline. You can use an RPC or indexed-data API. Or you can obtain prebuilt blockchain datasets for the specific protocols you need.
datastore.sh focuses on the third approach. Historical Solana datasets are delivered as documented, typed, and versioned Parquet files. The catalog includes protocol-specific data for areas such as Pump.fun, PumpSwap, Jupiter, Raydium, Meteora, Drift, Kamino, Orca, and Solana core programs.
The objective is not to replace every RPC use case. It is to remove the indexing and decoding work between:
“I need to research what happened on this Solana protocol”
and:
“I can run the query.”
Frequently asked questions about historical Solana data
Where can I download historical Solana data?
Historical Solana data can come from archival RPC infrastructure, blockchain-data APIs, data warehouses, or downloadable datasets. datastore.sh provides protocol-specific historical Solana datasets as partitioned Parquet files that can be downloaded and queried in your own infrastructure.
What is a Solana data provider?
A Solana data provider supplies access to blockchain information generated by the Solana network. Depending on the provider, this can include RPC access, transactions, account state, parsed instructions, protocol events, APIs, database tables, or bulk historical files.
Is Solana blockchain data public?
Solana’s ledger records public blockchain activity, but converting ledger history into convenient analytical datasets still requires infrastructure, indexing, and—in the case of protocol-specific analysis—decoding.
What is the difference between a Solana RPC and a historical dataset?
RPC exposes methods for interacting with and retrieving information from Solana. A historical dataset is organized specifically for large-scale analysis and may contain previously collected and decoded blockchain history.
Can I use historical Solana data for backtesting?
Yes. Historical trades, instructions, and protocol events can be used to reconstruct past market conditions and test strategies. Backtests must still enforce strict time ordering so future information is not accidentally used to make historical decisions.
Can I get Pump.fun historical data?
Yes. Pump.fun activity can be reconstructed from Solana blockchain history or obtained as protocol-decoded tables. datastore.sh provides historical Pump.fun instructions and events as a dedicated Solana dataset.
Can I get PumpSwap historical data?
Yes. PumpSwap is a separate Solana program and can be represented as its own historical dataset containing PumpSwap-specific instructions and events. datastore.sh provides PumpSwap historical data as typed Parquet tables.
What format is best for large historical blockchain datasets?
For analytical workloads, columnar formats such as Parquet are commonly useful because they work with analytical engines and warehouses while allowing queries to read only the relevant columns.
The important distinction
When choosing a Solana data provider, do not ask only:
“Do you have Solana transactions?”
Ask:
“Can this dataset answer the question I am actually researching?”
For protocol analytics, the difference between raw transactions and decoded historical data can represent weeks or months of indexing, decoder development, backfills, and validation.
If you need historical Pump.fun, PumpSwap, Jupiter, Raydium, Meteora, or other Solana protocol data, browse the documented Solana datasets and inspect the available schemas before deciding which historical coverage you need.
Datasets in this post
Dataset · SolanaPump.fun45 tables · schema v1.0 · from $200Dataset · SolanaPump.fun Swaps32 tables · schema v1.0 · from $200Dataset · SolanaJupiter Swap21 tables · schema v1.0 · from $200Dataset · SolanaRaydium AMM V49 tables · schema v1.0 · from $200Dataset · SolanaMeteora DLMM73 tables · schema v1.0 · from $200Dataset · SolanaDrift V29 tables · schema v1.0 · from $200Dataset · SolanaKamino Lending45 tables · schema v1.0 · from $200Dataset · SolanaOrca Whirlpool54 tables · schema v1.0 · from $200Data Platform
datastore.sh engineering
The team that operates capture, decoding, and reconciliation for every dataset in the catalog.