Introduction
Canary Historian is a strong system of record for industrial time-series data. It captures high-speed values from pumps, tanks, production systems, and plant-floor applications while preserving tag metadata and quality.
Databricks is where many teams want to combine that history with maintenance records, production events, inspection results, and business context. SignalX connects the two by reading Canary views, tags, metadata, current values, recorded values, interpolated values, and summaries into organized cloud-ready datasets.

Once the Canary data is cleaned, formatted, and contextualized, SignalX uses Zerobus to write the prepared operational data directly to a Delta Table. This provides a direct path for making historian data available in Databricks for analytics, reporting, and further processing without relying on traditional batch ETL processes.
Challenge
Canary data is naturally tag-centric. Databricks workloads often need a consistent output shape. A pump model, for example, may need inlet pressure, outlet pressure, flow, mass flow, and status delivered as row records, a pivoted tag table, or an asset-search table.
Canary Views and Virtual Views help organize tags, but cloud pipelines still need repeatable rules for selecting views, carrying metadata and quality forward, and choosing raw versus aggregate reads. Raw reads preserve historian timestamps; interpolated or summary reads are the better fit when a shared interval grid is required.
Custom exports may work for a first use case, but they become harder to maintain as sites, assets, and tags grow.

Solution
SignalX acts as a bridge between Canary Historian and the Databricks Lakehouse. It connects through the Canary gRPC API, resolves the selected tags or assets, and converts the results into Apache Arrow records for downstream delivery.

Pipelines can be configured around inline tags, Canary SearchTags expressions, CSV tag lists, or asset search. SignalX then handles metadata lookups, batching, current reads, scheduled history micro-batches, optional pivoted tag history, asset-search output, and fixed-range history.
This lets operations teams keep their historian structure intact while giving data teams a cleaner model in Databricks.

Key Features
Connector Modes
- Tag browse: explicitly specify the Canary tag list to read.
- Tag Search: Search for matching Canary tags before reading data.
- Asset Search: Search for assets and read the selected asset attributes.
Supported Data
- Read current tag values.
- Read raw, recorded, interpolated, or summary of tag history.
- Optionally pivot tag history into columns for a wider Databricks-friendly output.
- Read raw, recorded, interpolated, or summary of asset history.
- Read data from Canary Views.
Delivery to Databricks with Zerobus
- Carry metadata, quality, timestamps, numeric values, and string values into downstream records.
- Write directly to a Databricks Delta Table through Zerobus.
- Reduce infrastructure requirements by eliminating the need to deploy or manage additional ingestion services such as Azure Event Hubs, Kafka brokers, or ADLS landing zones.
- Lower ingestion costs by removing the associated costs of external messaging, broker, or landing-zone services.
- Reduce configuration effort because no additional Databricks jobs are required to ingest data from Event Hubs, Kafka, or ADLS.
- Lower operational overhead by reducing the number of moving parts and avoiding ongoing maintenance of external messaging or storage services.
- Support faster data delivery into Delta tables by keeping data in its native binary, compressed format throughout the pipeline.
SignalX Connectivity Beyond Databricks

Although this article focuses on Canary Historian data for Databricks, SignalX is not limited to Databricks as an endpoint. It can deliver operational data to different systems depending on the architecture, including connectors such as Azure Event Hubs, Kafka, ADLS Gen2, S3, SQL databases, and Azure IoT Hub. This makes SignalX useful for projects that need to move historians, SCADA, industrial applications, or cloud data into different downstream platforms without building a separate custom pipeline for every use case.
SignalX also provides broader integration capabilities for industrial environments. Its WebApp, Orchestrator, and Worker Service architecture can be deployed across separate servers and network boundaries, helping proxy data between facility networks, on-premises systems, and cloud platforms. Jobs can support streaming, backfill, event-driven, and server-style patterns, while optional transformations help prepare data before delivery. For PI write-back scenarios, SignalX can also work with PI Buffer Manager to support buffered writes to PI collectives.
We can help
SignalX is designed to reduce custom integration work, simplify historian-to-cloud delivery, and make high-quality OT data available for analytics, reporting, Databricks Delta Table, and AI initiatives.
If you are planning a Canary Historian to Databricks project, SignalX can help you move from tag exports to analytics-ready operational datasets with less manual effort. At MetaFactor, we help industrial teams modernize operational data pipelines without disrupting the systems that already work. SignalX gives organizations a practical way to connect Canary Historian data to Databricks while preserving the structure and context that engineers depend on. Contact us

