Contents
- First-party data sources as the foundation of AI personalisation
- Resolving user identity without cookies
- Using a Customer Data Platform to integrate data
- Deterministic and predictive segments in marketing
- Recommendation systems and generated content personalisation
- Personalising a website and app for better conversion
- Optimising advertising campaigns with AI
- The future of AI personalisation in the context of regulation and privacy
Share
Personalisation with AI in online marketing works best when it is based on first-party data and a consistent user identity, rather than external cookies. After the limitation of 3rd-party cookies, web and app events, CRM data and transaction history are becoming more important, as they make it possible to tailor the offer, content and communication channel. At the same time, companies have to solve the problem of “who the user is” across different devices, so they can build one customer view and make more accurate decisions. In this section, we show which first-party data sources are the most valuable and how to approach identity resolution in practice while respecting privacy. You will also learn which tools and methods help connect signals into a “single customer view” and where limitations most often appear. This makes it easier to plan personalisation that is operationally feasible and ready for the cookieless world.
First-party data sources as the foundation of AI personalisation
The most valuable data for AI personalisation after moving away from 3rd-party cookies is first-party data: web/app events, CRM data, transaction history and responses to communication. In practice, this means combining analytical events (e.g. GA4/Firebase) with customer information in CRM (e.g. Salesforce/HubSpot) and purchases (e.g. Shopify/Stripe). These four data streams make it possible to build content and offer matching based on real behaviour, not assumptions. It is also important to take account of signals from emails and notifications (email/push), because they show what the user actually responds to.
Personalisation becomes more useful when first-party data is combined into scenarios that solve specific business problems, for example reducing returns. For instance, a fashion store can combine the “view_item” and “add_to_cart” events with return history to recommend sizes and styles with a lower return risk. This approach uses both intent (viewing and adding to basket) and outcome (return), which improves the quality of recommendations. The better the data describes the context of the purchase decision and its consequences, the more accurate the personalisation.
The foundation of effectiveness is data quality, because even large volumes can lead to “odd” recommendations when key fields are missing or duplicates occur. The most common problems are incorrect product mapping (SKU), duplicate users and a lack of context, for example no price, category or margin. In practice, stream validation is implemented (e.g. dbt tests, Great Expectations) and quality thresholds are set, for example blocking personalisation when more than 2% of events have null in the product_id field. This ensures models and personalisation rules operate on data that can be considered reliable for decision-making.
- Web/app events (GA4/Firebase) as a record of intent and behaviour
- CRM data (Salesforce/HubSpot) as context for the relationship and communication
- Transaction history (Shopify/Stripe) as a source of value and recognisable purchasing patterns
- Responses to communication (email/push) as a signal of preferred channel and content type
Resolving user identity without cookies
Resolving identity without cookies means recognising the same user across multiple devices using first-party identifiers and controlled signal matching. The most common methods are login, magic links and first-party identifiers (e.g. customer_id), which make it possible to connect activity into one continuous history. Where there is no hard identifier, probabilistic signal matching is used (time, device, behaviour patterns), taking privacy constraints into account. The aim is to obtain as consistent a customer view as possible without relying on 3rd-party cookies.
In practice, identity resolution supports the creation of a “single customer view”, i.e. one customer profile that combines events, transactions and marketing interactions. Tools such as Segment, mParticle or Tealium help with this by providing profile merging rules and a standardised approach to collecting and joining data. Whether personalisation will be consistent across channels and devices is determined by the merging rules, not by the number of sources alone. This makes it possible to reduce issues such as “two profiles for the same person” and increase the relevance of personalisation decisions.
Consent and preferences also play a key role, and they should work as control data in user identification and activation. Marketing consent defines which channels and communication types are allowed (e.g. email, SMS, profiling), so it affects what the system can do with the profile and predictions. A good practice is to store consents as versioned records (timestamp, source, scope), so campaigns and models respect the current state and the history of changes. Without integrating consent into the personalisation logic, even well-resolved identity will not translate into correct marketing actions.
Using a Customer Data Platform to integrate data
A Customer Data Platform (CDP) helps integrate data by organising events into one common schema and making them available for near-real-time personalisation. In practice, a CDP standardises the data language, for example the “Purchase” event with properties such as value and currency, so different systems “understand” the same signals. This makes it easier to activate data in multiple places at once, without having to rewrite the logic for each tool separately. The greatest value of a CDP is the consistency of events and their rapid distribution to channels, models and measurement.
A CDP is particularly useful when a company simultaneously uses CRM, analytics and automation tools, but needs one consistent event stream for personalisation decisions. For example, Segment can send the same “Checkout Started” event at the same time to Braze (automation), BigQuery (model) and Meta CAPI (attribution). As a result, marketing, analytics and modelling are based on the same event definitions, which reduces the risk of differing interpretations of results. This makes it easier to link personalisation in channels with how it is measured and then optimised.
Deterministic and predictive segments in marketing
Deterministic segments work well when you need a simple, easy-to-justify audience split, whereas predictive segments are useful when you want to forecast the next behaviours and control budget and contact frequency. A deterministic segment (e.g. “bought within 30 days”) is quick to implement and transparent, but it does not suggest what the user will do next. A predictive segment (e.g. propensity-to-buy > 0.7) allows decisions to be made on the basis of probability, rather than history alone. In practice, predictive segments are better at supporting optimisation of activities with a large offer and a short decision cycle.
Predictive segments are most often built on features derived from behaviour and transactions, such as RFM (recency, frequency, monetary), category browsing intensity, price sensitivity and “time to next purchase”. In propensity-to-buy or churn models (e.g. XGBoost, LightGBM), the number of returns, a drop in activity and responses to campaigns are also taken into account. Operational thresholds are then defined, e.g. churn_risk > 0.8 triggers a retention sequence, and propensity > 0.6 raises the bid in a remarketing campaign. This approach makes it possible to differentiate actions not only by “who was active”, but also by “who is most likely to respond now”.
Recommendation systems and generated content personalisation
Recommendation systems and generated content personalisation make it possible to tailor what the user sees and reads to their behaviour, context and business objective. In e-commerce, personalisation most often takes the form of product ranking, which takes into account not only preferences, but also whether you are optimising revenue, margin or turnover. Building such a ranking uses, among other things, collaborative filtering, matrix factorisation, sequential models (Transformers) and Learning-to-Rank. In practice, “similar products” are only the starting point — the biggest difference is made by a ranking aligned both to the user and to business priorities.
Generated content personalisation (NLG) works when the texts are selected based on intent and verified in experiments, rather than just “sounding nice”. AI can prepare multiple variants of headlines or descriptions for the same category, and the system selects the version based on the predicted CTR for a specific user and device. For example, ChatGPT/Claude can generate 5 headline variants for the “running” category, and then the variant is selected algorithmically. The effectiveness of NLG depends on A/B tests, because attractive language does not have to translate into conversion.
Personalising a website and app for better conversion
Personalising a website and app improves conversion when you change key interface elements without rebuilding the entire site. The hero banner, category order, product card recommendations and checkout content are most often personalised, e.g. a message about free delivery above a threshold. Such changes are particularly useful because they affect the “here and now” experience at the moment of browsing the offer and closing the basket. The most practical implementations start with the places with the highest influence on the decision: listing, product card and checkout.
You should implement on-site/in-app personalisation so that you can measure its impact on business results in parallel and iterate variants in a controlled way. Tools such as Optimizely, Dynamic Yield or VWO make it possible to launch variants and measure their impact on conversion and AOV. This allows you to assess whether personalised modules genuinely improve purchase propensity, rather than relying solely on clicks. The key is to treat personalisation as an experimentation system: you deploy variants, measure the effect and only then scale.
Optimising advertising campaigns with AI
Optimising advertising campaigns with AI means tailoring creatives, selecting products from the catalogue and better controlling conversion signals sent to platforms such as Meta and Google. In practice, AI supports dynamic creative optimisation (DCO), selection of products for delivery and optimisation of conversion events (e.g. purchase value) reported via CAPI. For example, a product feed can be enriched with margin and availability, and then the delivery of low-stock products can be limited so as not to waste budget on items that are unavailable. The most “business-oriented” approach is to optimise not only for purchase, but also for margin and real availability in the catalogue.
AI also helps maintain targeting effectiveness despite privacy constraints, provided you build lookalikes on your own data and send only aggregates to the platforms via Conversions API/Enhanced Conversions. For example, you can create a “high LTV 180 days” segment and use it to generate similar audiences, then assess quality by comparing CAC and ROAS between lookalike 1% vs 3%. This approach combines audience personalisation with hard validation of campaign economics. At the same time, it reduces dependence on mechanisms based on 3rd-party tracking.
The future of AI personalisation in the context of regulation and privacy
The future of AI personalisation will be shaped increasingly strongly by GDPR, tracking limitations and growing documentation and supervisory requirements for AI systems. Personalisation can be “profiling” within the meaning of GDPR, especially when you automatically assess preferences and target offers, so you need an appropriate legal basis and transparency of actions. In practice, this means clear messages, a processing activities register and objection and consent withdrawal mechanisms available in 1–2 clicks. At the same time, it is worth following the data minimisation principle: collect only the features necessary for the purpose and shorten the retention of behavioural events (e.g. 90–180 days), and instead of precise GPS location use an approximation to the city or delivery zone.
Security and ethics will become as important as model “accuracy”, because personalisation brings together multiple sources into a single user profile. Standard practice includes encryption at rest and in transit, RBAC/ABAC access control (e.g. in Snowflake/BigQuery), as well as access audits with alerts when someone exports large volumes of customer data. There is also the risk of discrimination if the model is trained on historical data and indirectly uses sensitive attributes (e.g. postcode as a proxy for income), which is why bias testing, constraints built into the model (fairness constraints) and policies that limit risky personalisation scenarios are used. In practice, “compliant” personalisation is one that can be defended on every front: technically (audit), legally (transparency) and socially (no discrimination).
From a technology perspective, there is a clear trend towards moving personalisation into owned channels and solutions that limit data exposure, because attribution and remarketing are becoming less precise (Privacy Sandbox, SKAdNetwork), while aggregates, modelling and server-side tracking (CAPI, server-side tagging) are gaining importance. On-device personalisation and federated learning approaches are being used more and more often, where the model learns locally and the server collects only anonymised updates. This reduces the risk of privacy breaches, although it can be harder to maintain and requires robust quality control of updates. At the regulatory layer, practices such as a model card, data documentation, bias testing and “human-in-the-loop” procedures for sensitive scenarios are becoming more important. Over the next 12–24 months, personalisation will also move towards agents that plan and execute campaigns (brief → creative assets → tests → optimisation) within guardrails and budgets, for example by proposing 3 experiments per week, running them on 10% of traffic and scaling only those with positive uplift and stable margin.
FAQ
Frequently asked questions
Which 1st-party data is most important in AI personalisation?
The most valuable sources are web and app events, CRM data, transaction history, and responses to email and push. These sources make it possible to tailor offers and content to the user’s real behaviour.
Can AI personalisation work without 3rd-party cookies?
Yes, if it is based on your own data and a consistent user identity. The article highlights that after cookies are restricted, the importance of 1st-party data and identity resolution grows.
How is user identity resolved across different devices?
Logins, magic links and first-party identifiers such as customer_id are most commonly used. When there is no hard identifier, probabilistic matching of signals is used with privacy in mind.
Why is a Customer Data Platform important for personalisation?
A CDP organises events into one common schema and makes them available for personalisation almost in real time. As a result, marketing, analytics and models work with the same event definitions.
When is it better to use predictive rather than deterministic segments?
Predictive segments are better when you want to forecast next behaviours and control budget or contact frequency. Deterministic segments work well for a simple, easy-to-justify audience split.
Which website or app elements are most often personalised to improve conversion?
Most often, the hero banner, category order, recommendations on product cards and checkout content are personalised. The article also points out that it is worth starting with listings, product cards and the basket.





