Cortex XSIAM · Ingestion Economics
XSIAM & Third-Party EDR
Telemetry Economics
How much data a customer's existing EDR pushes into XSIAM, what that costs against the daily ingestion commit, and why native Cortex agent telemetry changes the math. Built to start the conversation, not end it.
Figures are planning ranges assembled from vendor documentation and field reporting, not quotes. Every one of them moves with agent version, policy, filtering, and the workstation-to-server mix. Baseline against the customer's own metrics before anything reaches a proposal.
Interactive
Daily Ingest Estimator
Pick the EDR the customer runs today and the license tier they hold, enter the estate, and see what shows up in the XSIAM daily ingestion commit. The native Cortex XDR agent is always the fixed comparison, and every source is scored side by side at the bottom.
- Endpoint telemetry into the Analytics tier—
- Into the Cortex Data Lake tier—
- Not ingested into XSIAM—
- Kept out of the Analytics tier by this routing—
- Other committed log sources—
- Endpoints counted—
- Average per endpoint, per day—
- Retained for 30 days—
- Same estate on native Cortex XDR agents—
- Removed from the bill by displacing the agent—
- Older-agent premium in this estate—
- Against the 100 GB/day Analytics minimum—
- Analytics on this combination—
Endpoint counts drive the endpoint telemetry line only. The Analytics tier carries a documented 100 GB/day minimum, and the optional Cortex Data Lake tier adds a 50 GB/day minimum on top of that, per the Cortex XSIAM license tier documentation. This estimator sizes data volume. It does not price anything.
Every number above is a planning range. Replace it with the tenant’s own.
Do this before any proposalThis estimator exists to frame the conversation, not to size a deal. The moment you have tenant access, stop quoting these bands and measure. It takes three queries, none of which filter on a vendor name, so the tenant’s own labeling appears in the output instead of being guessed at.
- What is arriving, and how much — every source ranked by GB per day, with the customer’s own
_vendorand_productstrings printed alongside. - Is it growing, and did something change — a daily series that proves the first number is representative rather than a bad day.
- Where it is concentrated — the collectors and reporting devices producing the most data, which is usually where the actionable finding is.
Apples to Apples
Which License Are You Comparing?
Every volume number on this page means something different depending on the tier the customer holds. Name the tier before you compare, because the tier decides whether endpoint telemetry is metered at all — and whether the customer even owns an agent to compare against.
Yes, Cortex Analytics is available without the Cortex XDR agent
This comes up constantly, so state it plainly: NG-SIEM is itself an Analytics subscription tier. Palo Alto documents it as “an Analytics subscription tier that includes data collection and full automation, suitable for users who want to enhance their security without immediately replacing their existing SIEM and endpoint solutions,” with Core Analytics listed as included in the license alongside UEBA, correlation and alerting, threat hunting, and automation. A customer on NG-SIEM with no Cortex agent still gets analytics.
The two real constraints- Detector coverage, not licensing, is the limit. Analytics, IOC, and BIOC issues are only raised on normalized logs, while Correlation Rule issues are raised on normalized and non-normalized logs. On third-party EDR telemetry only a subset of analytics detectors function, because there are no on-endpoint causality chains to feed the strongest behavioral models. The licence says yes; the data says partially.
- NG-SIEM does not entitle an agent. In the base products table, Enterprise Runtime Security (XDR) is listed as an Add-on at NG-SIEM and Included in license at Enterprise and Premium, and the agent entitlement language — “entitles you to one Cortex XDR agent per endpoint” — appears at Enterprise, with no equivalent stated for NG-SIEM. So the “native agent telemetry is free” argument is not available to an NG-SIEM customer until they add that entitlement, which is exactly the conversation the estimator is built to open.
- Do not confuse the two “Runtime Security” products. They sit in different columns of the same table and get mixed up constantly. Enterprise Runtime Security (XDR) is the endpoint layer: “comprehensive endpoint and server protection by combining AI-driven analytics, endpoint controls, next-generation antivirus, and automated investigation.” Cloud Runtime Security is the Cortex Cloud layer: Cloud Detection and Response, Cloud Workload Protection, and Web Application and API Security, priced per workload with a minimum workload count. Cloud Runtime Security is an add-on at NG-SIEM and at Enterprise, and only included at Premium. If someone pushes back that “runtime security is a Cortex Cloud thing,” they are thinking of the second one.
Tier comparison for sizing conversations
| Tier | Core Analytics & UEBA | Cortex XDR agent entitlement | Third-party EDR telemetry | Who it is for |
|---|---|---|---|---|
| XSIAM NG-SIEM | Included | Add-onVia Enterprise Runtime Security (XDR) — the endpoint layer, not Cloud Runtime Security | Native ingestionCrowdStrike, Microsoft Defender for Endpoint, SentinelOne — and fully metered | Keeping the existing SIEM and endpoint tool for now |
| XSIAM Enterprise | Included | One per endpointPlus Host Insights and Extended Threat Hunting Data | Still meteredRunning both means paying for endpoint data twice | Replacing endpoint and SIEM together |
| XSIAM Premium | Included | IncludedEndpoints, and hosts or cloud workloads per subscription | Still metered | Enterprise scope plus cloud posture, cloud runtime, XTI, TIM, and ASM |
| XSIAM Enterprise Plus (legacy) | Included | RetainedExisting holders keep Enterprise features including cloud agent features | Still metered | Existing holders; the full cloud posture bundle now requires Premium |
Commit floors apply to every tier above: the Analytics tier carries a documented 100 GB/day minimum, and the optional Cortex Data Lake tier adds a 50 GB/day minimum once the mandatory Analytics tier license is met. That floor matters in the estimator — an estate that models out below 100 GB/day is still billed to the minimum, so reduction work below the floor buys headroom rather than savings. “Pro per GB” also appears in the license plan documentation covering XDR endpoint, network, identity, and forensics data along with cloud posture and cloud runtime data; confirm with the account team whether it is in play before using it in a comparison.
How to use this with the estimator- Set the tier first, then the vendor. The estimator's second control now carries the tier, and the footnote under the verdict changes with it, including a warning when “native Cortex XDR agent” is selected under NG-SIEM, which is the most common way these comparisons quietly stop being apples to apples.
- Compare like for like in a displacement pitch. Third-party EDR on NG-SIEM versus native agents on Enterprise is a tier change and an agent change. Say both out loud, or the volume delta will look like a discount the customer never agreed to.
- Check the result against the floor. The estimator now shows how far the modeled total sits above or below the 100 GB/day Analytics minimum.
Source Data
What Each EDR Actually Sends
Three vendors, three completely different shapes of data. Each card links to the captured sizing research behind it. The volumes are not comparable to each other, and none of them are comparable to what a native agent produces, because the vendors measure and filter differently.
SentinelOne Deep Visibility Depends entirely on the filter
SentinelOne streams endpoint events out through Cloud Funnel to an S3 bucket, which XSIAM then reads. There is no single “SentinelOne number” to quote. The Cloud Funnel export filter is the whole ballgame, and it moves the volume by about two orders of magnitude. Quote the wrong end of that range and you are off by 100x in either direction.
File Creation, File Modification, File Deletion, File Rename — cuts the export by two orders of magnitude, described in one case as the difference between more than 100 GB and less than 100 MB of daily ingest. Quoting 6 or 7 TB for 10,000 endpoints overstates it for almost every real deployment and will cost you credibility in the room.
| Cloud Funnel configuration | Per workstation-class agent / day | 10,000 endpoints, Windows-dominant | Confidence |
|---|---|---|---|
| Vendor baseline — all fields, no filter | ~180 MB (v22.3+), ~280 MB (v22.2 or older and macOS) | ~2.8–5 TB / day | Captured vendor-metric research |
| Field-typical — noisy classes trimmed, most detail kept | ~90–120 MB | ~1–1.5 TB / day | Field-reported band, not vendor-published |
| Heavy filtering — file events stripped at source (unvalidated) | single-digit MB | tens of GB / day | Inferred from a reported ~100x reduction. Not a documented target state |
| Outlier — unfiltered export on a chatty estate | ~600–700 MB | ~6–7 TB / day | Single practitioner report; upper bound only |
Linux changes the arithmetic completely and is modeled separately in the estimator: roughly 1 GB per Linux workstation and up to 8 GB per Linux server per day at baseline, from about 500,000 and 4,000,000 events per agent per day respectively. A few hundred Linux servers can outweigh every workstation in the estate, so an unverified server count is the fastest way to be wrong by a multiple in either direction. Per-OS and per-version figures come from the captured sizing research and are directional planning numbers, not vendor commitments.
What to watch- Ask for the filter, not the endpoint count. The Cloud Funnel query and field selection determine the volume far more than how many endpoints the customer has. Without it you are guessing across a 100x range.
- Ask for the agent version spread and the server count. These are the two inputs a customer can pull from their own console, and both move the total materially: an older Windows agent costs roughly 55% more per day than a current one for identical coverage, and servers run an order of magnitude above workstations.
- Establish compressed or uncompressed. XSIAM accepts these logs gzipped or uncompressed, and JSON compresses roughly tenfold. A number quoted without that qualifier is not usable.
- Every raw gigabyte that does arrive is charged. In an NG-SIEM deployment where the customer keeps SentinelOne, all ingested EDR logs count against the XSIAM base subscription daily GB/day allowance.
- Filtering trades cost for detection, and heavy filtering is not a supported plan. The file events that dominate volume are also what ransomware and staging behavior look like. The Cortex setup procedure instructs that all fields be included, and Palo Alto warns that external EDR filters can restrict the data analytics require — so any material reduction needs a per-detection dependency review before it is committed, which is scoped assessment work rather than a console setting. That is the honest version of the coexistence conversation, and it is also the services conversation.
- Replacing the agent removes the line item entirely. Move those endpoints to native Cortex XDR agents and the telemetry stops costing anything against the daily allowance, with no filtering tradeoff to negotiate. Note the tier dependency: the per-endpoint agent entitlement starts at XSIAM Enterprise, and is an add-on under NG-SIEM.
- SentinelOne Log Reduction & Alert Tuning Guide — the documented collection paths and their volume controls, XQL to baseline real ingest and troubleshoot the feed, issue exclusions and BIOC exceptions, parsing-rule practice, a nine-step run order, and an honest verdict on whether Cribl helps.
- Captured 10,000-endpoint sizing research — the per-OS and per-agent-version figures that now drive the estimator, preserved verbatim with an editorial note on how each one is used here.
Optimizing SentinelOne volume: the documented tier model Do this before filtering
Before anyone talks about dropping events, there is a supported answer to high-volume third-party EDR: decide where each dataset lands. Palo Alto documents three destinations with different costs and very different capabilities, and the choice is a configuration decision rather than a data-loss decision. The estimator above now models all three — use control 4, Where the endpoint data lands.
| Destination | Best for | AI/ML analytics | Correlation rules | Where data lives | Key restrictions |
|---|---|---|---|---|---|
| Analytics tier | High-value security logs needed for real-time AI/ML detection | Full support | Full support | Ingested and stored in XSIAM | None stated. Mandatory tier, 100 GB/day minimum |
| Cortex Data Lake tier | High volume, low real-time security value, needed for compliance, investigation or hunting | None | Full support, consumes CU | Ingested and stored in XSIAM | No XDM normalization, no out-of-the-box analytics, no stitching or enrichment; cannot store PANW firewall logs; optional add-on, 50 GB/day minimum once the 100 GB/day Analytics minimum is met |
| Federated Search | Massive historical archives, or data the customer wants to keep in their own cloud storage | None | Search only | Not ingested — stays in the customer’s S3, Azure Blob or GCS | Ad-hoc XQL investigation only: no correlation rules, scheduled queries, widgets or dashboards; slower queries; CSV, Parquet or JSONL in Hive structure; not enabled by default — Customer Support must turn it on |
Tier assignment is done exclusively through Parsing Rules: Settings → Configurations → Data Management → Parsing Rules, open the Both tab, right-click the [INGEST] rule for the dataset and choose Change Tier to Data Lake. XSIAM disables the default rule, writes a user-defined rule with tier=lake, and appends a _lake_raw suffix to the dataset name. You can switch tiers at any time, but the change applies only to data ingested afterwards — existing data is never moved between tiers.
- 1. Decide what you actually forward. The largest and fastest reduction is not a filter, it is the collection path: the SentinelOne event collector and v2 integration pull activities, threats and alerts without the Cloud Funnel raw telemetry feed at all. Defensible when the SOC does not hunt in raw EDR; a real capability loss if they do.
- 2. Tier what you do keep. Raw telemetry that exists for compliance, investigation or occasional hunting is exactly the documented profile for the Cortex Data Lake tier — high volume, low real-time security value.
- 3. Federate what is purely an archive. If the customer wants sovereignty or a long retention horizon, leave it in their own bucket and reach it with Federated Search.
- 4. Only then discuss filtering, and only with a per-detection dependency review behind it. Dropping events is the one lever on this list that loses data permanently.
- 5. Reprice the remainder. Displacing the third-party agent removes the endpoint line item entirely rather than relocating it, because native Cortex XDR agent telemetry is covered by the per-endpoint agent license and does not count against the daily allowance as of today.
- “The Data Lake tier is just cheaper storage” — not quite. It is significantly cheaper to ingest, but queries, dashboards, playbooks and public API calls consume Compute Units, any mixed-tier query touching a single Data Lake dataset incurs a full CU charge, and retention and event-forwarding costs follow the standard Analytics tier model. Cheaper is not free, and the CU model has to be in the conversation.
- “You keep everything, you just pay less” — no. The Data Lake tier has no XDM normalization, which for SentinelOne means the data does not land in
xdr_data. BIOCs and analytics detectors depend on that normalized dataset, so they lose it. What survives is XQL and correlation rules written against the raw dataset. Note also that a few sources currently receive analytics-like treatment in the Data Lake tier as a temporary configuration that is documented as planned for removal — do not design around it. - “80 to 95 percent reduction” — unverified. No vendor publishes that figure for this feed. The defensible numbers are the ones the customer measures: baseline with the ingestion-metrics queries first, then quote a reduction against their own starting point.
- “Native telemetry is subsidized” — say it precisely. The documented position is that native Cortex XDR agent endpoint and analytics telemetry does not count against the daily GB/day allowance because it is covered by the per-endpoint agent license, and that the per-endpoint agent entitlement starts at XSIAM Enterprise. “Subsidized” invites a pricing question nobody on the call can answer.
Full detail, including the XQL to baseline current volume before choosing a tier, is in the SentinelOne Log Reduction & Alert Tuning Guide. Source: Optimize data management in Cortex XSIAM and Configure Cortex Data Lake tier.
CrowdStrike Falcon Data Replicator Low volume, low fidelity
CrowdStrike's numbers look dramatically smaller, and that is the story rather than a footnote. FDR delivers compressed batches, and CrowdStrike simply provides far less raw event data than other vendors. Cheap to ingest, and correspondingly thinner to detect and investigate on.
| Endpoint type | Minimum compressed data / host / day | Total minimum at 10,000 |
|---|---|---|
| Windows | 2.5 MB | ~25 GB / day |
| macOS | 2.5 MB | ~25 GB / day |
| Linux | 10 MB | ~100 GB / day |
These are absolute minimums, published as averages of compressed data, and they are what the estimator uses for the published-minimum band — 2.5 MB per Windows or macOS host and 10 MB per Linux host, reproducing ~25 GB and ~100 GB per day at 10,000 hosts. Some sources give the Linux figure as an 8–10 MB range. Real environments run higher, driven by host activity, the customer's Falcon module subscriptions, and which filtering policies are applied; field reports of 10–40 MB per endpoint per day are common, which is the estimator's activity-adjusted band. CrowdStrike publishes no upper figure and no per-agent-version split, so every band above the floor is derived.
What to watch- The data arrives partial. Raw FDR events lack much of the context other vendors provide natively, so XSIAM has to spend extra backend work stitching FDR records to other event sources to make them usable for analytics and investigation.
- Analytics coverage is a subset. With no on-endpoint causality chains, only a portion of XSIAM's analytics detectors light up, primarily network, file, process, and registry. The most advanced behavioral detectors are not fully operational on this telemetry.
- Still charged. Exactly as with SentinelOne, if the customer keeps CrowdStrike and streams via FDR, that volume counts against the XSIAM base subscription daily allowance.
- Compression cuts both ways. Compressed-at-rest figures are not the same as parsed and indexed volume. Confirm which measurement a customer is quoting you.
Deeper background: the captured CrowdStrike sizing research preserves the original FDR minimums and sizing considerations verbatim.
Microsoft Defender for Endpoint Two very different numbers
Microsoft publishes no flat expected megabytes per day. They estimate individual log entries at roughly 500 bytes, with certain events spiking to 2.5 MB each. That leaves you sizing between two numbers that differ by twentyfold, and picking the wrong one is how proposals fall apart.
| Sizing basis | Per endpoint / day | Total at 10,000 | What you actually get |
|---|---|---|---|
| The "free" E5 baseline | 5 MB | ~50 GB / day | Heavily truncated, a subset of basic alerts, missing critical raw telemetry |
| Real incident response volume | 100 MB or more | ~1 TB / day | The full EDR visibility threat hunting and IR genuinely require |
The 5 MB per endpoint per day included with a Microsoft Sentinel E5 entitlement is where most Microsoft-led sizing conversations start, and it is what the estimator uses for the E5-included band — about 50 GB per day at 10,000 endpoints. Customers who have tried to run real investigations on it consistently report needing at least 100 MB per endpoint per day, which is the effective-IR band and lands at roughly 1 TB per day for the same estate. Microsoft publishes no flat expected megabytes per day and no per-OS split, so anything above that is derived.
What to watch- Server data is often simply absent. E5 and MDE telemetry frequently isolates or excludes server data unless Defender for Servers is licensed separately. When the customer says 10,000 endpoints, verify whether servers are in that number, because their event volume is far higher than workstations.
- Streaming path matters. Raw MDE events reach XSIAM through the dedicated Defender for Endpoint collector via Azure Event Hubs. The generic Event Hub collector is not suitable for EDR data because it does not support stitching.
- Analytics coverage is partial. XSIAM covers key network, file, process, and registry detections from MDE, but Microsoft cannot supply the patented on-endpoint causality chains a native Cortex agent produces, so the most advanced AI behavioral detectors are not fully operational on this data.
- The whole volume is charged. Keep Defender and stream raw logs into XSIAM through Graph API or Event Hubs, and that entire roughly 1 TB per day counts against the base subscription daily GB/day allowance.
Deeper background: the captured Microsoft Defender sizing research preserves the E5 baseline versus effective-volume framing and the cost-saving strategy verbatim.
Fidelity Comparison
Native Agent vs Third-Party EDR
Ingest volume is only half the conversation. The other half is what XSIAM can do once the data lands, and that gap is where the commercial argument actually lives.
What “no charge” actually means here
XSIAM is sold with a daily data allowance measured in gigabytes per day. That allowance is the meter, and it is the thing the customer pays for. Anything that lands in it consumes allowance the customer bought. Anything that does not land in it is free to send, no matter how much of it there is.
An endpoint agent produces an enormous amount of data. Which side of that meter it lands on is decided entirely by whose agent produced it, and that single fact is worth more on most deals than any discount you could negotiate.
Native Cortex XDR agent — no charge
The customer already pays per endpoint for the agent, and the data that agent produces is covered by that same license. It never touches the gigabytes-per-day allowance. Ten thousand endpoints or a hundred thousand, the endpoint data adds nothing to the ingest bill. You are sizing the agent count, not the data volume.
Third-party EDR — charged
SentinelOne, CrowdStrike, and Defender data arrives as ordinary logs from an outside system, so it is billed exactly like firewall or cloud logs. Every gigabyte eats into the daily allowance, which means the customer is paying twice: once to their EDR vendor for the agent, and again to bring its output into XSIAM.
One caveat worth stating out loud. This is the commercial treatment as of today, not a contractual guarantee, and it is the single most consequential assumption behind every number on this page. Confirm current terms with your Palo Alto Networks account team in writing before you put a figure in front of a customer.
| Capability | Native Cortex XDR agent | SentinelOne | CrowdStrike FDR | Defender for Endpoint |
|---|---|---|---|---|
| Cost of endpoint telemetry against the daily GB/day allowance | No chargeIncluded with the agent license. Volume does not move the bill. | ChargedEvery GB consumes the paid daily allowance. | ChargedEvery GB consumes the paid daily allowance. | ChargedEvery GB consumes the paid daily allowance. |
| Cost of the extra analytics detail behavioral detection needs | No chargeAnalytics telemetry is part of the same agent entitlement. | ChargedArrives as raw logs, billed like any other source. | ChargedArrives as raw logs, billed like any other source. | ChargedArrives as raw logs, billed like any other source. |
| Raw events queryable in XQL | Yes | Yes | Yes | Yes |
Mapped into xdr_data and the XDM model |
Native | Yes | Yes | Yes |
| Event stitching across sources | Full | Supported, less optimized | Extra backend work needed | Dedicated collector required |
| On-endpoint causality chains | Patented, native | Not provided | Not provided | Not provided |
| Analytics-based detector coverage | Complete | Subset | Subset: network, file, process, registry | Subset: network, file, process, registry |
| Rule-based detection: correlation, BIOC, IOC | Yes | Yes | Yes | Yes |
| Agent-based prevention and response actions | Yes | Stays in vendor console | Stays in vendor console | Stays in vendor console |
| Exposure to vendor rate limits and export filters | None | Cloud Funnel limits apply | FDR batching and filters apply | Event Hub throughput applies |
Third-party EDR support is documented capability, not a compromise position. Palo Alto Networks documents that third-party agents typically provide less data than native agents, are not optimized to the same degree for causality analysis and cloud-based analytics, and that external EDR rate limits and filters can restrict the data comprehensive analytics needs.
The Native Path
How XSIAM Works With Itself
Coexistence with a third-party EDR is a real, supported deployment. Understanding what the native path does differently is what makes the coexistence conversation honest.
One agent, entitled per endpoint
XSIAM Enterprise entitles one Cortex XDR agent per endpoint, and licensing is a straight one-to-one count: one active device consumes one license. The agent handles tailored endpoint data collection and third-party log collection at the same time. That telemetry and the analytics telemetry derived from it are not billed against the daily ingestion commit today, which is why the endpoint line disappears from the sizing exercise on this path.
One schema, whatever the source
Whether the events come from a native agent, SentinelOne, CrowdStrike, or Defender, XSIAM parses and maps the critical fields into the xdr_data dataset and the XDM model, while keeping the raw events in their own vendor dataset for XQL. Analysts write one query pattern across every supported EDR. Fidelity varies by source, but the query surface does not.
Causality is the differentiator
The native agent builds causality chains on the endpoint itself, before anything is shipped. That structure is what the most advanced behavioral detectors consume. No third-party EDR provides it, so ingesting their telemetry gets you a genuinely useful subset of analytics detectors rather than the full set. This is a data availability limit, not a licensing gate.
NG-SIEM is the coexistence lane
XSIAM NG-SIEM is the analytics subscription tier for customers who are not ready to replace both their SIEM and their endpoint tool at once. It includes data collection and full automation, with native support for third-party EDR telemetry ingestion. It is the right first step for a CrowdStrike or Defender account, and it is also the deployment where the ingest bill is highest, because all that third-party raw telemetry is charged against the daily allowance.
Start Here
Discovery Questions
This page exists to open a conversation. These are the questions that turn a telemetry argument into a scoped deployment.
- "Of your endpoint estate, how many are servers rather than workstations?"Servers and Linux hosts run materially higher than workstations on every EDR, and CrowdStrike's own published averages put Linux at three to four times Windows. A 10,000-endpoint number with an unknown server mix is not a sizeable estate. Ask for the split before you model anything.
- "What are you ingesting into your current SIEM today, in GB per day, by source?"Their own metrics beat every published range. It also surfaces what they have already given up on collecting because of cost, which is usually where the value conversation actually starts.
- "Can you show me your Cloud Funnel export query and field selection?"For SentinelOne this single answer is worth about 100x in volume, since excluding file events alone reportedly cuts the export by two orders of magnitude. For Defender the equivalent question is whether they are living on the 5 MB E5 entitlement or the 100 MB reality. The export configuration is not a detail, it is the estimate.
- "When an analyst investigates an endpoint alert today, how many consoles do they touch?"Moves the conversation off gigabytes and onto the causality and stitching gap, where the native agent argument is strongest and easiest to demonstrate.
- "Is your EDR renewal coterminous with your SIEM renewal?"Determines whether this is a coexistence motion on NG-SIEM now with a native agent conversation later, or a single consolidation decision this cycle.
- "Do you know which of your detections and playbooks depend on the raw EDR feed you are paying to ingest?"Most teams cannot answer this, which is why data reduction stalls. Mapping that dependency before touching ingest is what makes a reduction plan safe to execute.
- "Are your servers licensed for EDR telemetry export at all?"Specific to Microsoft accounts. E5 and MDE telemetry often excludes server data unless Defender for Servers is purchased separately, so the visibility gap may already exist and be invisible in their current tooling.
Plain Language
Terms Without the Jargon
Everything on this page, said the way you would say it to someone who signs the contract rather than operates the tool.
xdr_data, and XQLProvenance
Sources
Licensing and capability claims trace to Palo Alto Networks documentation. Volume figures trace to vendor documentation and field reporting, and are planning ranges rather than commitments.
- Cortex XSIAM license tiers and product licenses — Palo Alto Networks Cortex documentationAnalytics tier is GB/day-based with a 100 GB/day minimum; optional Cortex Data Lake tier adds a 50 GB/day minimum. XSIAM Enterprise “entitles you to one Cortex XDR agent per endpoint,” with no equivalent entitlement stated for NG-SIEM. The base products table lists Enterprise Runtime Security (XDR) — endpoint and server protection with next-generation antivirus and automated investigation — as an Add-on at NG-SIEM and Included in license at Enterprise and Premium, and separately lists Cloud Runtime Security (CDR, CWP, WAAS, per-workload) as an Add-on at NG-SIEM and Enterprise and Included in license at Premium. Core Analytics is Included in license at every tier.
- Understand the Cortex XSIAM license plan — Palo Alto Networks Cortex documentationBase layer includes data storage, ingestion, query and reporting; agent entitlements and event forwarding add-ons.
- Ingest raw EDR events from SentinelOne Deep Visibility — Palo Alto Networks Cortex documentationCloud Funnel to S3 path, the
sentinelone_deep_visibility_rawdataset, mapping intoxdr_dataand XDM, and the documented third-party agent limitations. - Ingest raw EDR events from Microsoft Defender for Endpoint — Palo Alto Networks Cortex documentationAzure Event Hubs path, the dedicated collector requirement, the
msft_defender_rawdataset, and the analytics subset limitation. - Cortex XSIAM NG-SIEM license plans — Palo Alto Networks Cortex documentationNG-SIEM as the analytics tier for customers not replacing SIEM and endpoint simultaneously, with native third-party EDR telemetry ingestion from CrowdStrike, Microsoft Defender for Endpoint and SentinelOne. Base products list Data Collection plus Automation and Core Analytics for NG-SIEM, with Enterprise Runtime Security (XDR) appearing from Enterprise upward, and a separate “Pro per GB” line covering XDR endpoint, network, identity and forensics data with cloud posture and cloud runtime data.
- Understand the Cortex XSIAM product licenses (NG-SIEM documentation) — Palo Alto Networks Cortex documentationNG-SIEM described as “an Analytics subscription tier that includes data collection and full automation,” including out-of-the-box analytics, detection, threat hunting, response, automation and UEBA; Enterprise adds one Cortex XDR agent per endpoint; Premium adds the cloud posture and runtime bundle.
- Visibility of logs and issues from external sources — Palo Alto Networks Cortex documentation“While Correlation Rules issues are generated on non-normalized and normalized logs, Analytics, IOC and BIOC issues are only generated on normalized logs.” This is the detector-coverage limit behind the licence-versus-fidelity distinction on this page.
- Optimize data management in Cortex XSIAM — Palo Alto Networks Cortex documentationThe three documented destinations and their trade-offs: Analytics tier (full AI/ML analytics and correlation), Cortex Data Lake tier (no AI/ML analytics, no XDM modeling, no out-of-the-box analytics, correlation supported but consuming Compute Units), and Federated Search (not ingested, search only, Customer Support must enable it).
- Configure Cortex Data Lake tier — Palo Alto Networks Cortex documentation100 GB/day Analytics minimum and 50 GB/day Data Lake minimum; the capability matrix showing no XDM normalization, detections, stitching or enrichment in the Data Lake tier; Compute Unit charging for dashboards, playbooks, public API and queries, with full CU charge on mixed-tier queries; tier assignment through Parsing Rules with
tier=lakeand the_lake_rawdataset suffix; and the rule that tier changes are not retroactive. - License allocation — Palo Alto Networks Cortex documentationOne active device consumes one endpoint license, on a strict 1:1 basis.
- What's Next in Cortex: New Innovations for Security Operations — Palo Alto Networks blogGeneral availability of third-party EDR support, bringing XSIAM analytics to traditional endpoint tools.
- CrowdStrike FDR data volumes — Splunk Add-on for CrowdStrike FDR documentationAverage compressed data per host per day: 2.5 MB Windows, 2.5 MB macOS, 8–10 MB Linux, citing CrowdStrike.
- Falcon Data Replicator data sheet — CrowdStrikeFDR export mechanism and near real-time compressed batch delivery.
- SentinelOne Cloud Funnel ingestion filtering and best practices — Google Cloud Security CommunityPractitioner report that a default open Cloud Funnel query returns roughly 600–700 MB per endpoint per day, and that excluding File Creation, Modification, Deletion and Rename events reduces the export by two orders of magnitude, characterised as the difference between more than 100 GB and less than 100 MB of daily ingest.
- Cloud Funnel log format requirements — Palo Alto Networks Cortex documentationLogs must be one record per line and may be gzip compressed or uncompressed, which is why any quoted volume figure has to state which of the two it measures.
- The SentinelOne per-OS and per-agent-version baseline used by the estimator — roughly 180 MB for a v22.3+ Windows agent, 280 MB for a v22.2 or older Windows or macOS agent, about 1 GB per Linux workstation and up to 8 GB per Linux server, from about 90,000 / 140,000 / 500,000 / 4,000,000 events per agent per day — comes from the captured sizing research rather than a published vendor table.It reproduces the documented 10,000-endpoint totals of ~1.8 TB, ~2.8 TB, ~10 TB and ~80 TB per day for single-OS estates, and a 2.8–5 TB per day range for a standard Windows estate. The Microsoft Defender 5 MB entitlement versus roughly 100 MB effective volume, and the field-typical band that puts a tuned 10,000-endpoint estate near 1 TB per day, are field-reported planning ranges rather than vendor-published figures. All of it is directional: baseline against the customer's own export metrics or a proof of value before use in a proposal.