Migration Hadoop data to Databricks for ‘Gumtree’

Migration Hadoop data to Databricks for ‘Gumtree’

Migrating Hadoop data to Databricks for Gumtree

Gumtree DWH migration from legacy ebay systems to a scalable Lakehouse solution powered by Databricks and Google Cloud.

18 TB

Data Volume

5

Data Sources

25

Team Size

6 Months

Pipeline Build

Challenge

The Business Challenge

Gumtree needed to modernize its existing data warehouse environment by migrating data from legacy eBay systems to a cloud-based lakehouse architecture. The existing environment posed challenges in terms of high operational costs, complex infrastructure management, diverse data formats, scalability limitations, and the need for real-time data sharing with external vendors and enterprise applications.
The objective was to determine whether this process could be automated while maintaining accounting logic, mathematical accuracy and traceability to the underlying transactions.
icon1

Legacy Data Infrastructure

Existing data workloads were dependent on legacy Hadoop-based infrastructure, creating challenges around maintenance, scalability and modernization.

icon2

High Operational Costs

Maintaining hardware and IT infrastructure increased operational overhead and limited the ability to optimize resources dynamically.
icon

Diverse Data Formats

Needed to handle structured and semi-structured data across multiple formats. (JSON, Avro, ORC, Parquet, XML).
icon4

Scalability & Workload Isolation

Different application workloads required greater isolation and scalability to support growing data and processing requirements.
share

Real-Time Data Sharing

Business required the ability to share data in real-time with external vendors and other enterprise applications

Solution

A Modern Lakehouse on Databricks and Google Cloud

We designed and implemented a scalable lakehouse architecture on Databricks, hosted on Google Cloud, to migrate data from eBay systems. The solution leverages Databricks data pipelines and Parquet support to accelerate migration, enable workload isolation, and support multiple data formats.
Solution for Migration to Hadoop

Impact

Business Impact

The migration enabled Gumtree to move from a legacy Hadoop environment to a modern, scalable cloud-based lakehouse architecture.

Faster Data Delivery

Databricks data pipelines took less than 6 months to build, helping accelerate the migration.

Reduced Operational Costs

Lower hardware maintenance and IT infrastructure costs through Databricks cloud architecture.

Support for Multiple Data Types

Store and process structured and semi-structured data (JSON, Avro, ORC, Parquet, XML).

Greater Scalability

Application-level workload isolation provides more scalability and efficient resource utilization.

Real-Time Data Sharing

Easily share data in real time with external vendors and other enterprise applications within the organization.

TECHNOLOGY

Technology Approach 

Google Cloud

Cloud infrastructure foundation

Databricks

Lakehouse platform

Databricks Data Pipelines

Ingestion, transformation, and processing

Parquet

Optimized storage format

Enterprise Integrations

APIs, PostgreSQL, Hive, Salesforce, S3, GCS

Ready to Modernize Your Data Platform?

Move from legacy systems to a scalable lakehouse with Databricks and Google Cloud.

From GL Data to CFO Insight: Automating Financial Variance Analysis for Manufacturing

From GL Data to CFO Insight: Automating Financial Variance Analysis for Manufacturing

From GL Data to CFO Insight: Automating Financial Variance Analysis for Manufacturing

Turning complex transaction-level financial data into explainable, CFO-ready insights through intelligent automation.

Industry

Manufacturing

Function

Finance

Engagement

POC / Feasiblity

Solution

Financial Data Automation

From GL Data to CFO Insight: Automating Financial Variance Analysis for Manufacturing

Turning complex transaction-level financial data into explainable, CFO-ready insights through intelligent automation.

Industry

Manufacturing

Function

Finance

Engagement

POC / Feasiblity

Solution

Financial Data Automation

THE BUSINESS CHALLENGE

The client’s finance team spent significant time manually analyzing General Ledger movements, reconciling balances and preparing monthly variance commentary. With thousands of transactions, reversals, accruals and provisions contributing to a single GL movement, the process was complex, time-consuming and prone to inconsistency.

The objective was to determine whether this process could be automated while maintaining accounting logic, mathematical accuracy and traceability to the underlying transactions.

Large volume of transactions with reversals, accruals and provisions

Different accounting treatments for Balance Sheet and P&L accounts

Multiple line items contributing to one GL movement

Need for business explanations, not just numerical differences
High manual effort and long month-end cycle

THE SOLUTION

We built an intelligent GL Variance & Commentary automation framework that combines financial rules, transaction-level analysis, text normalization and intelligent commentary generation.

DATA SOURCE

Trial Balance

  • Current & comparison balances
  • GL descriptions
  • Overall GL variance

Transaction Dump

  • Transaction level details
  • Drivers behind the variance
  • Reversals, accruals, provisions

HOW IT WORKS

Ingest Financial Data

Upload Trial Balance and transaction dump along with base & comparison months.

Understand & Classify

Apply accounting logic for Balance Sheet (GL 1–4) and P&L (GL 5–6) accounts.

Normalize & Identify

Detect reversals, accruals, provisions and normalize text to extract meaningful drivers.

Calculate Movement

Aggregate transactions by driver and calculate month-over-month movements.

Generate Commentary

Create analyst-level and CFO-level commentary with quantified drivers and explanations.

Deliver Insights

View in dashboard and download Analyst & CFO commentaries (Excel).

From Numbers to Business Explanation

Example: GL 40000021

Driver Aug Accrual (LC) Sep Accrual (LC) Movement (LC)
Freight OL -2.63M -3.86M +1.23M
Engg OL -4.25M -4.54M +0.29M
IT OL -0.06M -0.29M +0.23M
Other movements — — -0.40M
Total +1.34M

The calculated movement reconciled exactly to the Trial Balance variance.

Analyst Commentary (Detailed)

The balance increased by $1.34M primarily due to higher accruals in Freight OL (+$1.23M) and Engineering OL (+$0.29M), partially offset by lower IT OL accruals (+$0.23M) and other offsetting movements (-$0.40M).

CFO Commentary (Business Summary)

Accruals increased by $1.34M mainly due to higher freight and engineering accruals, partially offset by lower IT accruals.

RESULTS FROM POC

11/11

GLs Tested

100%

Achieved ≥80% Variance Coverage

100%

Achieved ≥90% Variance Coverage

100%

Mathematical Accuracy Verified

80–90%+

Target Variance Explanation Coverage Achieved

BUSSINESS IMPACT

Reduced Manual Analysis

Automates the initial investigation of GL movements, reducing manual transaction review.

Faster Variance Review

Moves from manual analysis to a structured, repeatable variance workflow.

Greater Explainability

Quantifies drivers and explains why the variance occurred, not just that it changed.

Better Traceability

End-to-end traceability from Trial Balance → Drivers → Movement → Commentary.

CFO-Ready Insights

Two levels of output: Analyst detail and concise, business-focused CFO summary.

TECHNOLOGY APPROACH

Financial Data Processing

Large volume of transactions with reversals, accruals and provisions

Intelligent Data Processing

Text normalization, classification, reversal detection, accrual/provision handling

Analytics & Visualization

GL variance analysis, driver reconciliation, coverage measurement, interactive dashboard

AI Enhancement Layer

LLM-powered commentary generation with prompt tuning & cost optimization

Ready to Automate Your Financial Reporting?

Let’s turn your financial data into explainable insights that save time and drive better decisions.

AML Compliance Implementation for a US Financial Institution Using Apache Spark on Databricks Lakehouse

AML Compliance Implementation for a US Financial Institution Using Apache Spark on Databricks Lakehouse

AML implementation for US financial client

Anti Money Laundering Compliance implementation for US financial institutions using Apache Spark on Lakehouse

10 TB

Data Volume

4+

Data Sources

15

Team Size

3

Months to Compliance

AML implementation for US financial client

Anti Money Laundering Compliance implementation for US financial institutions using Apache Spark on Lakehouse

10 TB

Data Volume

4+

Data Sources

15

Team Size

3

Months to Compliance

1. THE CHALLENGE

Financial institutions face increasing pressure to detect suspicious activities, comply with stringent AML regulations, and manage large volumes of data from multiple sources. Traditional systems struggled to deliver real-time visibility and efficient analysis.

Complex Data Landscape

Large volumes of data from multiple heterogeneous sources.

Compliance Pressure

Need for faster detection and adherence to AML regulations.

Operational Inefficiencies

High infrastructure costs and limited scalability with legacy systems.

Risk of Fraud

Difficulty in identifying connected networks of suspicious individuals.

1. THE CHALLENGE

Financial institutions face increasing pressure to detect suspicious activities, comply with stringent AML regulations, and manage large volumes of data from multiple sources. Traditional systems struggled to deliver real-time visibility and efficient analysis.

Complex Data Landscape

Large volumes of data from multiple heterogeneous sources.

Compliance Pressure

Need for faster detection and adherence to AML regulations.

Operational Inefficiencies

High infrastructure costs and limited scalability with legacy systems.

Risk of Fraud

Difficulty in identifying connected networks of suspicious individuals.

2. THE SOLUTION

We implemented an AML solution using modern data technologies and a scalable Lakehouse architecture to enable real-time analytics and compliance.

Key Details

Dataset

10 TB of Data

Sources

  • Oracle
  • Postgres
  • Hive
  • HDFS from hadoop system

Team Size (15)

  • PM – 1
  • Tech Leads – 1
  • Data Engineers & Testers – 10
  • Managed Service Support Engineers – 3

Solution Highlights

  • Built unified data pipelines on Apache Spark with Lakehouse architecture.
  • Integrated multiple data sources to create a single source of truth.
  • Implemented GraphFrames to identify connected networks of suspicious individuals.
  • Accelerated AML compliance processes and reporting.
  • Leveraged Databricks cloud-based platform for scalability and cost efficiency.

3. THE RESULT

We implemented an AML solution using modern data technologies and a scalable Lakehouse architecture to enable real-time analytics and compliance.

We used GraphFrames in implementation that are helpful to identify the connected network of suspicious individuals transacting illegally

US Finance compliance on AML met within 3 months of project launch.

Reduce operational costs associated with hardware maintenance and IT infrastructure management through Databricks cloud-based architecture, optimizing resources and budget allocation.

3. THE RESULT

We implemented an AML solution using modern data technologies and a scalable Lakehouse architecture to enable real-time analytics and compliance.

We used GraphFrames in implementation that are helpful to identify the connected network of suspicious individuals transacting illegally
US Finance compliance on AML met within 3 months of project launch.

Reduce operational costs associated with hardware maintenance and IT infrastructure management through Databricks cloud-based architecture, optimizing resources and budget allocation.

Reimagining Sales Efficiency with Snowflake 

Reimagining Sales Efficiency with Snowflake 

Overview

A globally recognized real estate conglomerate — known for its extensive portfolio of luxury and commercial properties — set out to modernize how its sales teams accessed and interacted with business data. With sales executives distributed across geographies, the leadership identified inefficiencies in data retrieval, lead management, and reporting workflows. They envisioned an AI-powered internal sales assistant that could deliver instant insights, reduce manual effort, and drive faster decision-making.

KloudPortal was engaged as the strategic technology partner to design and deploy this solution, leveraging Generative AI, Snowflake, and Microsoft Copilot for seamless enterprise integration.

The Challenge

The sales and marketing functions were hindered by fragmented data residing across CRMs, property management systems, and reporting dashboards. Sales executives had to manually prepare reports or depend on data teams for basic queries, leading to delays in follow-ups and missed revenue opportunities.

The organization’s goals were clear:

  • Unify data sources into a single access layer.
  • Build a conversational AI that could query real-time data in natural language.
  • Integrate the assistant into daily workflows via Microsoft Copilot.
  • Ensure enterprise-level compliance, security, and scalability.

The Solution

KloudPortal deployed a multidisciplinary team of Snowflake, Generative AI, and Cloud Engineers to design a future-ready architecture centered around data intelligence and AI orchestration.

1. Data Centralization with Snowflake

The foundation began with consolidating all sales, lead, and marketing data within Snowflake’s Data Cloud. Using Snowpipe and Streams, KloudPortal engineered near real-time ingestion from Salesforce, ERP, and analytics systems. This created a single source of truth accessible through secure Snowflake APIs.

2. Generative AI Intelligence Layer

The team then implemented a Generative AI layer powered by OpenAI’s GPT models and LangChain orchestration. The chatbot was trained with domain-specific vocabulary — property details, lead statuses, sales KPIs, and performance metrics — to respond with precision and context.

It could now handle complex multi-turn conversations such as:

“Show me the top-performing projects this quarter.”
“What’s the conversion trend by region?”
“Which deals need immediate follow-up this week?”

3. Architecture & Deployment

Sales Efficiency with Snowflake
The solution followed a modular, cloud-native architecture for scalability and compliance:

  • Data Layer: Snowflake (AWS) for centralized, structured storage.
  • AI Layer: OpenAI GPT models managed via Azure OpenAI Service.
  • Integration Layer: LangChain middleware to route and translate natural language queries into Snowflake SQL commands.
  • Access Layer: Microsoft Copilot integration, allowing sales teams to query the chatbot directly from familiar Microsoft 365 applications such as Outlook and Excel.
  • Security & Compliance: Snowflake RBAC policies, encryption-in-transit, and Azure AD-based identity management ensured data privacy and governance.

4. Cloud Automation & CI/CD

The deployment leveraged Azure DevOps pipelines, enabling automated build, test, and release processes. Infrastructure-as-Code (IaC) principles ensured consistent provisioning and environment scalability.

The Impact

Within six months of implementation, the organization achieved transformative results:

  • 70% reduction in data retrieval and reporting time.
  • 25% improvement in sales funnel velocity.
  • Real-time decision-making through AI-based insights integrated directly into Copilot.
  • Increased productivity as sales teams focused more on clients and less on data handling.

The AI assistant evolved into a trusted digital co-worker, driving operational intelligence across departments.

5-Year ROI Projection

Over five years, the AI-powered sales assistant delivered a remarkable ROI by optimizing sales efficiency and decision-making. With over 14,000 hours saved annually, productivity gains exceeded $3.2 million, while improved lead conversions generated an additional $4.5 million in revenue. Combined with reduced IT and reporting overheads, the solution created a total value impact of approximately $9 million. This transformation demonstrates how data-driven intelligence can directly translate into measurable business growth and sustainable operational efficiency.

Reimagining Sales Efficiency with Snowflake
Metric Annual Impact 5-Year Projection
Time Saved by Sales Teams (~14,000 hours/year) $640,000 in productivity gain $3.2M
Improved Lead Conversion (12% → 15%) ~$900,000 added revenue annually $4.5M
Reduced IT & Reporting Costs ~$260,000 annual savings $1.3M
Total 5-Year ROI — ~$9 Million

Technologies Used

  • Snowflake – Unified data layer for analytics and sales operations.
  • OpenAI GPT & LangChain – Core Generative AI and orchestration engine.
  • Azure Cloud – Model hosting, integration, and security infrastructure.
  • Microsoft Copilot – Conversational interface integrated into sales workflows.
  • Azure DevOps – Continuous deployment and environment automation.

Conclusion

By combining Snowflake’s unified data architecture with Generative AI intelligence and Copilot integration, KloudPortal delivered a solution that redefined sales enablement for one of the world’s largest real estate enterprises.

The AI-powered assistant transformed how data was accessed and acted upon — enabling a culture of agility, accuracy, and intelligent automation. With a projected ROI of nearly $9 million over five years, this initiative stands as a benchmark for AI-driven sales transformation in enterprise real estate.

Case Study:   Building an Intelligent, Scalable Forecasting & Supply Chain Platform  

Case Study:   Building an Intelligent, Scalable Forecasting & Supply Chain Platform  

Executive Summary

Global supply chains today face unprecedented volatility, with challenges ranging from shifting consumer demand to supplier disruptions. Traditional forecasting engines, reliant on rigid ETL pipelines and monolithic models, are inadequate to handle this complexity. Our company partnered with a global enterprise to design a pioneering, modular, plugin-driven forecasting and planning platform. By decomposing forecasting into reusable business components—forecasting, disaggregation, allocation, validation, and optimization—we delivered a next-generation planning ecosystem. This platform drives precision forecasting, agile scenario planning, and scalable execution across millions of SKUs and multiple geographies, empowered by scalable cloud-native technologies.

Client Profile

A Fortune 500 global enterprise with diversified product lines serving multiple regions faced operational risks and financial inefficiencies due to legacy planning systems unable to adapt to modern supply chain complexity.

Strategic Challenges

  • Data Chaos: Multiple ERP and supplier feeds with inconsistent rules created fragmented, unreliable data inputs.
  • Forecast Volatility: Wild swings in demand predictions due to noisy and unstandardized inputs hampered confident planning.
  • Inventory Imbalance: Overstock situations in some regions contrasted with crippling stock-outs elsewhere, increasing working capital burdens.
  • Lack of Scalability: Legacy ETL and forecasting engines couldn’t scale to accommodate growing SKU counts, product launches, and geographic expansion. These challenges threatened service levels, working capital efficiency, and ultimately, the client’s market share and operational resilience.

Our Approach: Modular Planning Architecture

We re-imagined forecasting not as a monolith but as a library of composable, domain-specific business plugins, orchestrated for efficiency and scale. This plugin-driven modularity allows rapid customization and agility in evolving market conditions.
  • Data Integrity & Governance: Validation plugins ensured a single source of truth by ingesting only clean, standardized data across brands, geographies, and horizons.
  • Forecasting & Disaggregation: Ensemble forecast plugins projected demand at national, regional, and SKU levels; disaggregation plugins translated forecasts into granular actionable demand signals.
  • Allocation & Optimization: Allocation modules prioritized inventory against supply constraints, regional demands, and strategic goals; optimization plugins balanced costs, service levels, and supply resilience.
  • Scalable Execution Framework: Plugins were categorized into HeavyWeight, MediumWeight, and PySpark types to orchestrate computing resources efficiently. This design supports processing millions of records daily, enabling near real-time decision-making.

Business Impact

  • Enhanced Forecast Accuracy: Improved prediction reliability by 25–30%, significantly reducing the risk of misinformed decisions.
  • Working Capital Efficiency: Freed up millions in capital via inventory reductions driven by precise allocation strategies.
  • Service Level Improvement: Decreased stock-outs by over 15%, boosting customer satisfaction and loyalty.
  • Agile Scenario Planning: Enabled rapid “what-if” simulations, shortening decision cycles by up to 50% during market disruptions.
  • Scalability & Future-Proofing: Delivered a platform that scales seamlessly with new products, markets, and data volumes, powered by cloud-native, event-driven architecture.

5 Years RoI Projection

5 Years RoI Projection
Over five years, the modular, cloud-native forecasting platform is expected to evolve from delivering immediate efficiency (forecast accuracy, inventory reduction) to becoming a strategic asset that drives enterprise-wide agility and resilience. The cumulative ROI is projected to climb from 150% in Year 1 to 1,700% by Year 5, solidifying it as one of the highest-value digital transformation initiatives for the client. Here is a projection summary of the RoI Year-on year, along with the key business impact it is going to make.
Year Key Business Impact Annual Savings ($M) Cumulative ROI (%) Notes
Year 0 Initial investment phase — 0% $2M investment in design, integration, and setup
Year 1 Enhanced forecast accuracy (25–30%) 3M 150% Rapid efficiency gains from better demand planning
Year 2 Reduced stock-outs & improved service 5M 350% Stabilized planning and cross-region optimization
Year 3 Working capital optimization 7M 650% Mature forecasting models and automation
Year 4 Advanced scenario planning & AI-driven optimization 9M 1000% Predictive simulations improve resilience & agility
Year 5 Continuous learning & enterprise-wide rollout 12M 1400% Platform scaled across global business units
Case Study: Python Upgrade & Data Modernization for a Global Supply Chain Company 

Case Study: Python Upgrade & Data Modernization for a Global Supply Chain Company 

Client Overview

A leading global Supply Chain Management Company that relies heavily on custom-built software products to streamline logistics, inventory, and procurement workflows. Their technology stack included Python-based applications and numerous custom plugins that powered core business processes.

Business Challenge

The client’s entire product base was running on Python 3.10 (2021 release), which was reaching the end of long-term support. Continuing with this version posed several risks:

  • Security Vulnerabilities: Lack of patches and updates would expose the system to threats.
  • Compatibility Issues: Third-party libraries and plugins were evolving rapidly, with some deprecating support for older Python versions.
  • Performance Bottlenecks: Legacy CSV-based data handling created inefficiencies in large-scale data processing and analytics.
  • Future Readiness: The system needed to be prepared for upcoming enhancements in data processing and cloud tenant management (BDP Tenant).

Solution Approach

Our team partnered with the client’s in-house developers to execute a phased upgrade and migration plan:

  1. Python Upgrade
    • Migrated the entire product suite from Python 3.10 (2021) to Python 3.11 (2023).
    • Identified and refactored custom plugins, internal libraries, and third-party dependencies to ensure full compatibility.
    • Implemented automated test coverage to validate business-critical workflows.
  2. Data Modernization
    • Transitioned from CSV-based data storage to Apache Parquet format for faster, more efficient, and scalable data processing.
    • Optimized ETL pipelines for analytics and reporting use cases, leveraging Parquet’s columnar format.
  3. Cloud & Multi-Tenant Enablement
    • Integrated with BDP (Big Data Platform) Tenant Management, enabling better multi-tenant data governance and scalability.
    • Designed flexible configurations for future product expansion across regions.
  4. Performance Optimization & Future Readiness
    • Benchmarked Python 3.11’s new performance enhancements (up to 10–60% faster execution for certain workloads).
    • Ensured the product architecture was future-proof to adopt upcoming Python releases and data innovations.

    Key Results

    • Seamless Migration: All core applications and plugins upgraded without downtime.
    • 30–40% Faster Processing: Achieved significant performance improvements due to Python 3.11 optimizations and Parquet adoption.
    • Improved Data Efficiency: Reduced storage footprint by 40% and accelerated analytics workloads.
    • Stronger Security & Compliance: Eliminated risks of outdated dependencies.
    • Future-Ready Platform: Positioned the client for scaling and adopting advanced analytics, AI, and multi-tenant architectures.

    Conclusion

    This migration not only secured the client’s current operations but also set the foundation for scalable innovation. By upgrading to Python 3.11 and modernizing data formats, the client now has a robust, efficient, and future-ready platform that supports supply chain excellence.

We use cookies to ensure that we give you the best experience on our website. If you continue to use this site we will assume that you are happy with it.

Cookies settings
Accept
Privacy & Cookie policy
Privacy & Cookies policy
Cookie name Active

Privacy Policy

What information do we collect?

We collect information from you when you register on our site or place an order. When ordering or registering on our site, as appropriate, you may be asked to enter your: name, e-mail address or mailing address.

What do we use your information for?

Any of the information we collect from you may be used in one of the following ways: To personalize your experience (your information helps us to better respond to your individual needs) To improve our website (we continually strive to improve our website offerings based on the information and feedback we receive from you) To improve customer service (your information helps us to more effectively respond to your customer service requests and support needs) To process transactions Your information, whether public or private, will not be sold, exchanged, transferred, or given to any other company for any reason whatsoever, without your consent, other than for the express purpose of delivering the purchased product or service requested. To administer a contest, promotion, survey or other site feature To send periodic emails The email address you provide for order processing, will only be used to send you information and updates pertaining to your order.

How do we protect your information?

We implement a variety of security measures to maintain the safety of your personal information when you place an order or enter, submit, or access your personal information. We offer the use of a secure server. All supplied sensitive/credit information is transmitted via Secure Socket Layer (SSL) technology and then encrypted into our Payment gateway providers database only to be accessible by those authorized with special access rights to such systems, and are required to?keep the information confidential. After a transaction, your private information (credit cards, social security numbers, financials, etc.) will not be kept on file for more than 60 days.

Do we use cookies?

Yes (Cookies are small files that a site or its service provider transfers to your computers hard drive through your Web browser (if you allow) that enables the sites or service providers systems to recognize your browser and capture and remember certain information We use cookies to help us remember and process the items in your shopping cart, understand and save your preferences for future visits, keep track of advertisements and compile aggregate data about site traffic and site interaction so that we can offer better site experiences and tools in the future. We may contract with third-party service providers to assist us in better understanding our site visitors. These service providers are not permitted to use the information collected on our behalf except to help us conduct and improve our business. If you prefer, you can choose to have your computer warn you each time a cookie is being sent, or you can choose to turn off all cookies via your browser settings. Like most websites, if you turn your cookies off, some of our services may not function properly. However, you can still place orders by contacting customer service. Google Analytics We use Google Analytics on our sites for anonymous reporting of site usage and for advertising on the site. If you would like to opt-out of Google Analytics monitoring your behaviour on our sites please use this link (https://tools.google.com/dlpage/gaoptout/)

Do we disclose any information to outside parties?

We do not sell, trade, or otherwise transfer to outside parties your personally identifiable information. This does not include trusted third parties who assist us in operating our website, conducting our business, or servicing you, so long as those parties agree to keep this information confidential. We may also release your information when we believe release is appropriate to comply with the law, enforce our site policies, or protect ours or others rights, property, or safety. However, non-personally identifiable visitor information may be provided to other parties for marketing, advertising, or other uses.

Registration

The minimum information we need to register you is your name, email address and a password. We will ask you more questions for different services, including sales promotions. Unless we say otherwise, you have to answer all the registration questions. We may also ask some other, voluntary questions during registration for certain services (for example, professional networks) so we can gain a clearer understanding of who you are. This also allows us to personalise services for you. To assist us in our marketing, in addition to the data that you provide to us if you register, we may also obtain data from trusted third parties to help us understand what you might be interested in. This ‘profiling’ information is produced from a variety of sources, including publicly available data (such as the electoral roll) or from sources such as surveys and polls where you have given your permission for your data to be shared. You can choose not to have such data shared with the Guardian from these sources by logging into your account and changing the settings in the privacy section. After you have registered, and with your permission, we may send you emails we think may interest you. Newsletters may be personalised based on what you have been reading on theguardian.com. At any time you can decide not to receive these emails and will be able to ‘unsubscribe’. Logging in using social networking credentials If you log-in to our sites using a Facebook log-in, you are granting permission to Facebook to share your user details with us. This will include your name, email address, date of birth and location which will then be used to form a Guardian identity. You can also use your picture from Facebook as part of your profile. This will also allow us and Facebook to share your, networks, user ID and any other information you choose to share according to your Facebook account settings. If you remove the Guardian app from your Facebook settings, we will no longer have access to this information. If you log-in to our sites using a Google log-in, you grant permission to Google to share your user details with us. This will include your name, email address, date of birth, sex and location which we will then use to form a Guardian identity. You may use your picture from Google as part of your profile. This also allows us to share your networks, user ID and any other information you choose to share according to your Google account settings. If you remove the Guardian from your Google settings, we will no longer have access to this information. If you log-in to our sites using a twitter log-in, we receive your avatar (the small picture that appears next to your tweets) and twitter username.

Children’s Online Privacy Protection Act Compliance

We are in compliance with the requirements of COPPA (Childrens Online Privacy Protection Act), we do not collect any information from anyone under 13 years of age. Our website, products and services are all directed to people who are at least 13 years old or older.

Updating your personal information

We offer a ‘My details’ page (also known as Dashboard), where you can update your personal information at any time, and change your marketing preferences. You can get to this page from most pages on the site – simply click on the ‘My details’ link at the top of the screen when you are signed in.

Online Privacy Policy Only

This online privacy policy applies only to information collected through our website and not to information collected offline.

Your Consent

By using our site, you consent to our privacy policy.

Changes to our Privacy Policy

If we decide to change our privacy policy, we will post those changes on this page.
Save settings
Cookies settings