Galileo.ml

You didn't hire PhDs to clean data.

Galileo gives funds training-ready market data in 24 hours. Your team builds the models. We handle everything underneath them.

Time to first trained model
Before Galileo 12+ weeks
Clean
Format
Integrate
5 raw feeds Messy data Training
With Galileo 24 hours
Galileo
One simple integration Training
  • The highest data quality in the marketPoint-in-time correct, dual-sourced, validated on every run.
  • The fastest integration you will runOne call. Live in your own stack in 24 hours instead of months.
  • Nothing left to cleanDelivered in your schema, into your bucket, ready to train.

Hundreds of funds train on Galileo.

We help funds train their ML models faster. Every fund we talk to says the same thing: the people they hired to find alpha are stuck cleaning tickers. We take that off their desk and hand back data that is clean, point-in-time correct, and ready to train the day it lands.

We were three weeks into building our own ingestion layer when we signed. We deleted it.
Head of Research Systematic equity fund, $2.4bn AUM
Two of our best factors stopped working when we replayed them properly. Better to find that in research than in the book.
Portfolio Manager Multi-strategy fund, $6bn AUM
It lands in our bucket every morning in the schema we asked for. Nobody has raised a data ticket in four months.
Chief Technology Officer Quant credit fund, $900m AUM

What we cover.

Everything below is point-in-time and arrives in your own schema.

Global financial markets

Prices, OHLC and volumes across equities, ETFs and indices.

Fundamentals

Standardised line items, stored at the date they were filed.

Rates and employment

Policy rates, yield curves and labour releases, as they printed.

Oil, gas and commodities

Spot and forward curves, inventories and production figures.

Sentiment

Earnings calls, news, filings and social media scored per company and timestamped to the minute.

Alternative data

Consumer, web and supply-chain signals, mapped to the companies they move.

Custom datasets

Tell us what you need and we build it. If the dataset does not exist yet we source it, clean it and deliver it on the same terms as everything above.

One call and it's in your stack.

We deliver into whatever you already use. Nothing to install, and the data never leaves your own environment.

  Most data vendors Galileo
Setup Six to twelve weeks 24 hours
Delivery Their portal, or files you parse yourself Straight into your S3, Snowflake or API
History Back-filled and quietly overwritten Every vintage kept and addressable
Support Open a ticket and wait Fixed by a dedicated engineer
Your researchers Cleaning, formatting, chasing mistakes Model building, logic and research
Amazon S3Google Cloud StorageSnowflake DatabricksParquetArrow PythonRESTgRPC

Pick one and we write to it. If you switch later, it is a config change.

Get a data sample

Talk to us and tell us the universe, the fields and the years you need. Your team will have a data link to it the same day.