Skip to content
datastore.sh

Hyperliquid Historical Data: Trades, Funding, Order Books & More

Learn where Hyperliquid historical data comes from, what the API and public archives contain, their limitations, and how to get trades, funding, order books and ledger data for backtesting.

Data Platform16 min read

If you are looking for Hyperliquid historical data, there is an important detail to understand before downloading anything:

There is no single Hyperliquid endpoint that gives you every type of historical data in one clean table.

Hyperliquid exposes different pieces of history through its API, public S3 archives, and node-generated data. The official historical-data documentation separates those sources clearly.

Depending on what you are researching, you may need:

  • historical trades and fills
  • wallet-level fills
  • funding payments
  • funding rates
  • L2 order-book snapshots
  • liquidations
  • deposits and withdrawals
  • ledger changes
  • staking and delegation activity
  • validator rewards
  • block-level activity

Those datasets answer very different questions.

A trader backtesting an execution strategy needs something very different from an analyst studying wallet profitability or a researcher reconstructing Hyperliquid’s order book.

This guide explains what Hyperliquid historical data is available, where it comes from, the limitations of the official API and archives, and what you actually need for quantitative research.

What historical data is available on Hyperliquid?

Hyperliquid’s historical data can broadly be divided into four categories.

Trade and fill data

Trade data tells you what actually executed.

Typical fields include:

  • market
  • price
  • size
  • buy or sell side
  • timestamp
  • wallet
  • order ID
  • trade ID
  • maker/taker information
  • fees
  • realized PnL
  • starting position
  • liquidation information

This is normally the starting point for historical volume analysis, wallet analysis, strategy backtesting, PnL research, maker/taker studies, liquidation analysis, and market-impact research.

Hyperliquid’s node data contains fill records, and the official historical-data documentation points to the hl-mainnet-node-data S3 bucket for historical fills.

Current data is written under node_fills_by_block, while older periods use different historical prefixes. That distinction matters if you are attempting to construct a long continuous dataset yourself.

Funding data

Perpetual futures on Hyperliquid use funding payments between long and short positions.

Historical funding data can answer questions such as:

  • What was the funding rate for BTC at a specific time?
  • Which markets consistently traded with positive funding?
  • How much funding did a wallet pay?
  • How much funding did a wallet receive?
  • Does extreme funding predict later price movement?
  • What was the real carrying cost of a historical position?

For a realistic perpetuals backtest, funding cannot simply be ignored.

Consider a strategy that holds a position for three weeks. Its approximate trading result is not simply:

text
exit value - entry value

You may need:

text
trading PnL
- trading fees
+/- funding payments
= strategy PnL

A strategy that appears profitable using only entry and exit prices can become significantly less attractive after fees and funding are included.

Historical order-book data

Trade data tells you what executed. Order-book data tells you what liquidity was available before execution.

That distinction is critical for execution research.

A historical trade might tell you:

text
BTC traded at $100,000
size = 0.5 BTC

But that does not tell you whether a hypothetical 50 BTC market order could also have executed at $100,000.

To estimate that, you need historical market depth. An L2 order book generally groups resting liquidity by price level:

text
Bid price    Bid size
99,999       2.4 BTC
99,998       5.1 BTC
99,997       7.8 BTC

and:

text
Ask price    Ask size
100,001      1.7 BTC
100,002      3.9 BTC
100,003      8.2 BTC

With this data you can study bid/ask spread, market depth, order-book imbalance, liquidity around major events, estimated slippage, execution cost, and short-term price pressure.

Hyperliquid maintains historical L2 book snapshots in its public hyperliquid-archive S3 bucket. The official documentation says that this archive is uploaded approximately monthly, has no guarantee of timely updates, and may contain missing data.

That is an important caveat if your research requires continuous, verified order-book coverage.

Ledger and account events

Not all valuable Hyperliquid data is market data.

Hyperliquid also generates events associated with movements of assets and protocol state. Examples include:

  • deposits
  • withdrawals
  • transfers
  • staking deposits
  • delegations
  • staking withdrawals
  • validator rewards
  • funding distributions
  • other ledger updates

These records are useful when the research question is about users rather than markets.

For example:

How much capital has a particular wallet deposited?

or:

Which addresses receive the largest validator rewards?

or:

What happened to balances around a particular sequence of trades?

These questions cannot necessarily be answered correctly from a trade table alone.

Where can you get Hyperliquid historical data?

There are several ways to obtain it.

Hyperliquid’s API

Hyperliquid provides a public Info API containing many useful market and user endpoints. For small historical queries, it can be the easiest place to start.

For example, you can retrieve user fills using the userFillsByTime Info endpoint documented in the official API reference:

json
{
  "type": "userFillsByTime",
  "user": "0x...",
  "startTime": 1750000000000,
  "endTime": 1750100000000
}

But there are limits that become important for historical research. Hyperliquid currently documents userFillsByTime as returning at most 2,000 fills per response, with only the 10,000 most recent fills for the address available through that query.

Imagine researching a highly active market-making wallet. If that wallet has executed 500,000 fills, repeatedly querying its current API history does not automatically reconstruct its complete trading career.

The API is extremely useful. But API availability and complete historical coverage are not the same thing.

When the API works well

The API is a good choice when you need:

  • recent wallet fills
  • current market information
  • current positions
  • current order books
  • a limited historical range
  • application-level queries

When the API becomes awkward

Bulk research becomes harder when you need:

  • every trader
  • every market
  • billions of fills
  • long time ranges
  • repeated backtests
  • large cross-wallet studies

At that point, continuously paging an API is often the wrong data architecture.

Hyperliquid’s public historical S3 data

Hyperliquid also publishes historical information through AWS S3. According to its official documentation, historical sources include:

text
s3://hyperliquid-archive/

and:

text
s3://hl-mainnet-node-data/

The archive contains data such as L2 order-book snapshots and asset contexts. Node data contains historical fills, blocks, and L1 transactions.

One important operational detail is that the S3 buckets are requester-pays. The data itself is publicly accessible, but the requester pays the AWS transfer costs.

Downloading data therefore involves more than clicking a CSV file. You generally need to:

  1. identify the correct bucket;
  2. determine the correct prefix;
  3. enumerate the historical files;
  4. pay AWS transfer costs;
  5. download compressed files;
  6. decompress them;
  7. parse their schemas;
  8. normalize changing historical formats;
  9. validate coverage; and
  10. convert the result into an analytical format.

For a one-off engineering project, that may be acceptable. For a researcher who simply wants a table of historical fills, it can become a significant amount of infrastructure work.

Run a Hyperliquid node

Another approach is to run infrastructure yourself and collect data as Hyperliquid produces it.

Hyperliquid documents node datasets including transactions, trades, order statuses, raw book differences, and miscellaneous events. Its node data schema documentation notes that default node data generation can produce roughly 100 GB of logs per day.

Running a node is attractive when you need complete control over future collection, low-level data, custom indexing, real-time research, or your own historical archive going forward.

But a node that starts today does not magically provide every dataset from the past. Historical backfilling is a separate problem.

Use a prepared Hyperliquid historical dataset

The other option is to use a provider that has already converted historical Hyperliquid data into analytical tables.

This is the approach datastore.sh takes with its Hyperliquid Historical dataset.

Instead of starting with compressed archive objects and building your own decoding and normalization pipeline, you receive typed historical tables in Parquet.

The current dataset contains nine documented tables:

TableWhat it contains
swapsSpot and perpetual fills
fundingPer-user funding history
l2_order_book_snapshotsHistorical L2 market depth
ledger_updatesBalance and protocol ledger changes
depositsUser deposits
withdrawalsUser withdrawals
delegationsValidator delegation activity
validator_rewardsValidator rewards
gossip_auctionsGossip auction activity

The important point is that these are separate typed tables rather than one enormous generic event blob. That makes the data much easier to use for analytical workloads.

What should Hyperliquid historical trades contain?

Trade data seems simple until you begin doing real analysis.

At minimum you probably want:

  • timestamp
  • market
  • price
  • size
  • direction
  • wallet
  • order ID
  • trade ID

Advanced research may also require:

  • spot vs. perpetual
  • maker vs. taker
  • fee
  • fee token
  • realized PnL
  • position direction
  • starting position
  • TWAP status
  • liquidation status

Suppose you want to measure Hyperliquid’s daily trading volume. One potential mistake is counting both counterparties in every matched trade.

A trade has a buyer and a seller. If your dataset contains one fill row for each participant and you blindly sum both sides, you can effectively count market volume twice.

A proper dataset needs enough information to identify the relevant side or correctly deduplicate matches.

Hyperliquid liquidation data

Liquidations are particularly valuable for market-structure research. You might want to study liquidation volume by market, long vs. short liquidations, liquidation clusters, market movement immediately after liquidations, wallets repeatedly liquidated, liquidation price, position size at liquidation, or the relationship between liquidations and funding.

One useful fact about Hyperliquid’s data model is that liquidation information can be associated with fill records rather than existing only as an entirely separate market feed. Liquidation analysis can therefore start with fill history and filter the relevant liquidation context.

For example:

sql
SELECT
    block_time,
    coin,
    wallet,
    px,
    sz,
    notional_usd
FROM swaps
WHERE is_liquidation = 1
ORDER BY block_time;

From there you can aggregate liquidation notional over time or compare it with price movements.

Hyperliquid funding data for research

Funding is useful for more than accounting. Researchers commonly examine relationships between:

text
funding rate
open positioning
market direction
price momentum
liquidations

One possible analysis is to measure forward returns after unusually high funding:

sql
SELECT
    coin,
    event_time,
    funding_rate,
    funding_amount,
    szi
FROM funding
WHERE ABS(funding_rate) > threshold;

You could then join those observations to later trade prices.

But there is an important methodological problem: do not use future funding information when constructing a historical trading signal.

If a model makes a decision at 10:00:00, the features used by that model must have been available by 10:00:00. Otherwise you have created look-ahead bias.

Hyperliquid order-book data for backtesting

Order-book history becomes particularly important when a backtest needs realistic execution.

Consider this simplistic backtest:

text
Signal appears
BTC price = $100,000
Buy 100 BTC at $100,000

That assumes $10 million of liquidity existed at exactly that price. Maybe it did. Maybe the actual order book looked like this:

text
Ask                  Available BTC
$100,000             0.8
$100,010             2.1
$100,025             4.9
$100,050             8.7
...

A 100 BTC market order would move through many levels of the book. Your real average execution price could therefore be substantially different from the quoted market price.

This is why candlestick data is not enough for execution-quality backtesting. OHLC candles show open, high, low, and close; they do not show how much liquidity was available at each price.

Historical L2 data helps answer that question.

What can you research with Hyperliquid historical data?

Once the main datasets are joined correctly, you can investigate much more than simple price history.

Wallet profitability

For each wallet:

text
realized PnL
- fees
+/- funding
= approximate realized trading result

This can be segmented by market, strategy duration, maker/taker activity, long/short exposure, and time period.

Trader behavior

You can study how frequently a wallet trades, average position size, markets traded, holding behavior, maker/taker preference, liquidation frequency, realized PnL, and funding paid or earned.

Market microstructure

Using fills and order-book data, researchers can measure spread, depth, imbalance, slippage, price impact, liquidity recovery, and volatility following large trades.

Liquidation research

Combine trades with liquidation information to investigate liquidation cascades, liquidation concentration, price movement around forced closes, recurring liquidation levels, and long/short asymmetry.

Quantitative backtesting

Historical data can be used to test strategies involving funding, momentum, mean reversion, order-book imbalance, whale activity, liquidation events, market depth, and wallet behavior.

Why Parquet is useful for Hyperliquid data

Historical market data gets large quickly. For analytical workloads, Parquet has several advantages over JSON or CSV. It is columnar, compressed, typed, and efficient for scans, and it is supported by most analytical engines.

Suppose your fills dataset contains 37 columns but your research requires only:

text
block_time
coin
wallet
px
sz

A columnar engine can read the columns it needs rather than repeatedly processing every field.

You can query Parquet directly with DuckDB, Polars, ClickHouse, Spark, Snowflake, or BigQuery.

For example, with DuckDB:

sql
SELECT
    coin,
    SUM(notional_usd) AS volume
FROM read_parquet(
    'hyperliquid/historical/swaps/schema=v1.0/date=*/part-*.parquet'
)
WHERE is_taker = 1
GROUP BY coin
ORDER BY volume DESC;

No application API needs to sit between the query and the data.

API vs. downloadable Hyperliquid data

The right approach depends on what you are building.

Use caseBetter starting point
Current market stateAPI or WebSocket
Current positionsAPI
Recent fills for one walletAPI
Application backendAPI
Record new data going forwardNode or stream
Study millions of historical tradesBulk dataset
Analyze many walletsBulk dataset
Large quantitative backtestBulk dataset
Historical market-depth researchOrder-book archive or dataset
Repeated ML feature generationBulk dataset
Build your own complete indexerNode plus archives

There is no universal winner. The mistake is using an API designed for application queries as if it were automatically a complete historical research database.

How much Hyperliquid history is actually available?

This question needs to be answered per table, not per provider.

For example, datastore.sh currently documents its Hyperliquid Historical bundle as covering July 2025 through July 2026 overall. But the individual tables do not all begin on the same date.

Examples from the current dataset include:

  • swaps: July 2025 → July 2026
  • funding: September 2025 → July 2026
  • l2_order_book_snapshots: May 2026 → June 2026
  • gossip_auctions: April 2026 → July 2026

That distinction matters. If your research requires both historical trades and historical order books, your usable study period is constrained by their overlapping coverage.

A provider saying “We have one year of Hyperliquid data” is therefore not enough. Ask: “One year of which table?”

Questions to ask a Hyperliquid data provider

Before using a historical dataset for research or backtesting, verify:

  1. What exact date range does each table cover?
  2. Does trade data include both maker and taker records?
  3. How are duplicate trade sides identified?
  4. Are wallets included?
  5. Are fees included?
  6. Is realized PnL included?
  7. Are liquidations explicitly identifiable?
  8. Does funding contain actual wallet payments or only rates?
  9. Is the historical book L2, L3, or L4?
  10. How frequently is the order book sampled?
  11. Are spot and perpetual markets distinguished?
  12. Are timestamps exchange timestamps, block timestamps, or receipt timestamps?
  13. Are decimal quantities stored exactly?
  14. How are missing periods represented?
  15. Can you inspect sample data before buying?
  16. Is the schema versioned?
  17. Are files checksummed?
  18. Can you retain and query the data yourself?

These questions are more important than the headline number of rows.

Common Hyperliquid historical-data mistakes

Treating API history as unlimited history

A paginated endpoint can still have a total historical limit. Always verify both.

Counting both sides of a trade as separate market volume

Market-wide volume usually requires correctly identifying or deduplicating the matched trade.

Ignoring funding in perpetual backtests

Long-running perpetual strategies can accumulate meaningful funding payments.

Using candles to simulate large executions

Candles contain price summaries, not historical liquidity.

Assuming every dataset has the same historical range

Trade history and order-book history often have different coverage.

Using floating-point arithmetic for financial aggregation

Exact decimal fields should generally remain decimal while you aggregate them.

Accidentally introducing future information

Every feature in a historical trading decision must be restricted to information available at that moment.

Where can I download Hyperliquid historical data?

There are three practical routes.

Use Hyperliquid directly if you want to build your own pipeline from the API, public S3 archives, and node data.

Run a node if you want complete control over data collection going forward.

Use a prepared historical dataset if your objective is analysis rather than building the underlying indexing infrastructure.

datastore.sh provides its Hyperliquid Historical dataset as typed, partitioned Parquet files that can be downloaded and queried in your own environment.

The dataset currently contains nine documented tables covering trades/fills, funding, L2 order books, ledger updates, deposits, withdrawals, delegations, validator rewards, and gossip auctions.

Files are delivered with versioning, manifests, and SHA-256 checksums. You can inspect the schemas and samples before choosing the historical coverage you need.

Explore Hyperliquid historical data →

Frequently asked questions about Hyperliquid historical data

Does Hyperliquid provide historical data?

Yes. Hyperliquid provides historical information through its API as well as public AWS S3 archives and node data. The available history, format, and limitations vary by data type.

Where can I get Hyperliquid historical trades?

Historical fills are available through Hyperliquid’s node-data archives. Recent wallet fills are also available through the Hyperliquid Info API. For large-scale analysis, historical fills can also be obtained as prepared analytical datasets.

How far back does the Hyperliquid API go?

There is no single limit applying to every API method. Limits vary by endpoint. For example, Hyperliquid currently documents userFillsByTime as returning at most 2,000 fills per response, with only the 10,000 most recent fills available for a user through that method.

Can I download Hyperliquid order-book history?

Yes. Hyperliquid publishes historical L2 book snapshots in its public hyperliquid-archive S3 bucket. Hyperliquid notes that the archive is uploaded approximately monthly, is not guaranteed to update on time, and may contain missing data.

Can I get Hyperliquid liquidation data?

Yes. Liquidation information can be derived from Hyperliquid fill data containing liquidation context. Prepared datasets can expose this as explicit fields to make liquidation analysis easier.

Can I get historical Hyperliquid funding data?

Yes. Hyperliquid exposes funding-related information through its data infrastructure and API. Prepared historical datasets can also provide funding payments, funding rates, and associated position information in analytical tables.

Is Hyperliquid historical data free?

Hyperliquid publishes public historical data, but its historical-data documentation states that the requester pays AWS transfer costs. You may also incur infrastructure and engineering costs to download, normalize, store, and process the data.

What format is best for Hyperliquid historical data?

For large analytical workloads, Parquet is useful because it provides columnar storage, compression, typed fields, and compatibility with engines including DuckDB, Polars, ClickHouse, Spark, and major cloud warehouses.

Can I use Hyperliquid historical data for backtesting?

Yes. Historical fills, funding, and order-book data can support strategy backtesting. A robust backtest should also account for fees, funding, available liquidity, slippage, and strict event-time causality rather than relying only on historical candle prices.

Does datastore.sh provide Hyperliquid historical data?

Yes. datastore.sh currently provides a Hyperliquid Historical dataset consisting of nine typed tables delivered as partitioned Parquet. The documented dataset includes swaps, funding, L2 order-book snapshots, ledger updates, deposits, withdrawals, delegations, validator rewards, and gossip auctions.

The key takeaway

If you search for Hyperliquid historical data, the challenge is not simply obtaining a list of old prices.

The right dataset depends on the question.

For price research, trades may be enough.

For perpetual strategy analysis, add funding and fees.

For realistic execution backtests, add historical order-book depth.

For wallet research, you need wallet-level fills and ledger activity.

And if you need to rerun those analyses hundreds of times, owning a structured historical dataset can be considerably simpler than rebuilding the same history from APIs and compressed archives each time.

The best starting point is therefore not:

“How do I download all Hyperliquid data?”

It is:

“Which historical events do I need to answer my research question?”

Datasets in this post

Dataset · HyperliquidHyperliquid Historical9 tables · schema v1.0 · from $200

Data Platform

datastore.sh engineering

The team that operates capture, decoding, and reconciliation for every dataset in the catalog.