Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.
Big Data AnalyticsEasy
Q1. What is Apache Spark?
- A.A production web server for hosting applications
- B.A standalone relational database management system
- C.A general-purpose programming language runtime
- D.A fast distributed computing engine for big data processing✓ Correct
Explanation
Apache Spark is a unified analytics engine for large-scale data processing, known for its in-memory computing capabilities making it faster than MapReduce.
Report an error in this question
Big Data AnalyticsEasy
Q2. What is Hadoop?
- A.A standalone relational database for transactional queries
- B.A multi-layer deep neural network architecture for AI
- C.An open-source framework for distributed storage and processing of big data✓ Correct
- D.A general-purpose compiled programming language for applications
Explanation
Apache Hadoop is an open-source framework that allows distributed storage (HDFS) and processing (MapReduce) of large datasets across clusters.
Report an error in this question
Big Data AnalyticsEasy
Q3. Big Data is characterized by:
- A.Only the fast speed of arrival
- B.Only the diverse data types
- C.Volume, Velocity, Variety (3 Vs)✓ Correct
- D.Only the large volume of data
Explanation
Big Data is commonly defined by the 3 Vs: Volume (large amount), Velocity (fast generation), and Variety (diverse formats).
Report an error in this question
Big Data AnalyticsEasy
Q4. What is MapReduce?
- A.A comparison-based sorting algorithm for small arrays
- B.A programming model for processing large datasets in parallel across a cluster✓ Correct
- C.A structured query language for relational database management
- D.A supervised machine learning algorithm for classification
Explanation
MapReduce splits data processing into Map (parallel processing) and Reduce (aggregation) phases for distributed computation.
Report an error in this question
Big Data AnalyticsEasy
Q5. What is a data warehouse?
- A.A centralized repository for structured data used for analysis and reporting✓ Correct
- B.A multi-layer deep neural network trained on data
- C.A physical room for storing server hardware equipment
- D.A transactional database used for daily operational queries
Explanation
A data warehouse stores large volumes of structured, historical data from multiple sources, optimized for analytical queries and reporting.
Report an error in this question
Big Data AnalyticsEasy
Q6. What does ETL stand for?
- A.Edit, Transfer, Link
- B.Extract, Transform, Load✓ Correct
- C.Encrypt, Transmit, Log
- D.Evaluate, Test, Learn
Explanation
ETL is the process of Extracting data from sources, Transforming it into the desired format, and Loading it into a target system like a data warehouse.
Report an error in this question
Big Data AnalyticsEasy
Q7. What is HDFS?
- A.A general-purpose interpreted programming language runtime
- B.A desktop operating system for personal computers
- C.A multi-layer deep neural network for classification
- D.Hadoop Distributed File System for storing data across a cluster✓ Correct
Explanation
HDFS (Hadoop Distributed File System) stores large files across multiple machines in a cluster, providing high throughput access to data.
Report an error in this question
Big Data AnalyticsMedium
Q8. What is Apache Kafka?
- A.A distributed event streaming platform for real-time data pipelines✓ Correct
- B.A standalone relational database for transactional workloads
- C.A supervised machine learning framework for model training
- D.An interactive data visualization and charting dashboard
Explanation
Apache Kafka is a distributed streaming platform for building real-time data pipelines and streaming applications with high throughput and low latency.
Report an error in this question
Big Data AnalyticsEasy
Q9. What is distributed computing?
- A.Cloud storage of files without any computation tasks
- B.Spreading computation across multiple machines working together✓ Correct
- C.Using a single powerful dedicated computer for processing
- D.Interactive data visualization and charting tools
Explanation
Distributed computing divides tasks across multiple networked computers that work together, enabling processing of data too large for a single machine.
Report an error in this question
Big Data AnalyticsEasy
Q10. What is a data lake?
- A.A cleaned and normalized relational database table store
- B.An interactive dashboard for data visualization tools
- C.A storage repository holding raw data in its native format✓ Correct
- D.A small curated dataset for specific model training
Explanation
A data lake stores vast amounts of raw data in its native format (structured, semi-structured, unstructured) until needed for analysis.
Report an error in this question
Big Data AnalyticsMedium
Q11. What is stream processing?
- A.Only storing data without processing it
- B.Processing data in real-time as it arrives✓ Correct
- C.Processing data in large scheduled batches
- D.Only visualizing data without analysis
Explanation
Stream processing analyzes and acts on data continuously as it arrives, enabling real-time analytics and immediate responses to events.
Report an error in this question
Big Data AnalyticsEasy
Q12. What is batch processing?
- A.Processing data interactively with user input prompts
- B.Processing data manually without any automation
- C.Processing large volumes of data as a group at scheduled intervals✓ Correct
- D.Processing data in real-time as each record arrives
Explanation
Batch processing handles large accumulated data in groups at scheduled times, as opposed to real-time stream processing.
Report an error in this question
Big Data AnalyticsMedium
Q13. What is data partitioning?
- A.Encrypting data at rest for privacy and compliance
- B.Dividing data into smaller pieces distributed across multiple nodes✓ Correct
- C.Merging all distributed data together into a single node
- D.Deleting old data records to free up storage space
Explanation
Data partitioning distributes data across multiple storage nodes or processing units, enabling parallel processing and improving query performance.
Report an error in this question
Big Data AnalyticsMedium
Q14. What is data sharding?
- A.Encrypting data at rest for privacy regulation compliance
- B.Copying the entire database to every single server node
- C.Splitting a database into smaller pieces across multiple servers✓ Correct
- D.Deleting old data records to reclaim disk storage space
Explanation
Sharding horizontally partitions data across multiple database servers, each holding a subset of the data, enabling horizontal scaling.
Report an error in this question
Big Data AnalyticsMedium
Q15. What is the role of a scheduler in big data systems?
- A.Training supervised machine learning classification models
- B.Visualizing data trends using interactive chart dashboards
- C.Managing and coordinating the execution of distributed tasks✓ Correct
- D.Storing and persisting data on distributed file systems
Explanation
A scheduler (e.g., YARN, Mesos) manages resource allocation and coordinates the execution of distributed processing tasks across cluster nodes.
Report an error in this question
Big Data AnalyticsMedium
Q16. What is Apache Hive?
- A.A distributed event streaming platform for real-time pipelines
- B.A graph database for storing relationship network data
- C.A data warehouse tool providing SQL-like queries over Hadoop data✓ Correct
- D.A supervised machine learning library for model training
Explanation
Apache Hive provides a SQL-like interface (HiveQL) for querying and managing large datasets stored in Hadoop's distributed storage.
Report an error in this question
Big Data AnalyticsMedium
Q17. What is the CAP theorem?
- A.It is primarily about effective data visualization techniques
- B.It applies only to single standalone machines and servers✓ Correct
- C.All three properties can always be fully achieved simultaneously
- D.A distributed system can guarantee at most two of: Consistency, Availability, and Partition tolerance
Explanation
The CAP theorem states that a distributed system can provide at most two of three guarantees: Consistency, Availability, and Partition tolerance simultaneously.
Report an error in this question
Big Data AnalyticsMedium
Q18. What is NoSQL?
- A.An interactive data visualization and charting dashboard
- B.Non-relational databases designed for flexible schemas and horizontal scaling✓ Correct
- C.A general-purpose compiled programming language for apps
- D.A newer version of the standard SQL query language syntax
Explanation
NoSQL databases (e.g., MongoDB, Cassandra) provide flexible schemas and horizontal scalability for handling large volumes of diverse data types.
Report an error in this question
Big Data AnalyticsMedium
Q19. What is the difference between Spark and Hadoop MapReduce?
- A.They are completely identical technologies with no performance differences
- B.Spark only works effectively on small datasets under one gigabyte
- C.MapReduce is significantly faster than Spark for all data workloads
- D.Spark processes in-memory making it faster; MapReduce writes to disk between steps✓ Correct
Explanation
Spark performs in-memory computation, making it up to 100x faster than Hadoop MapReduce for certain workloads, which reads/writes to disk between map and reduce steps.
Report an error in this question
Big Data AnalyticsMedium
Q20. What is a Spark DataFrame?
- A.A simple NumPy array structure✓ Correct
- B.A distributed collection of data organized into named columns
- C.A Python pandas DataFrame structure only
- D.A traditional database table structure only
Explanation
A Spark DataFrame is a distributed collection of data organized into named columns, similar to a pandas DataFrame but distributed across a cluster for parallel processing.
Report an error in this question
Big Data AnalyticsHard
Q21. What is data skew in distributed processing?
- A.Perfectly even and balanced distribution of data across all nodes
- B.Uneven data distribution across partitions causing some nodes to be overloaded✓ Correct
- C.Data records that are missing or have null values throughout
- D.Data records that have been accidentally duplicated in storage
Explanation
Data skew occurs when data is unevenly distributed across partitions, causing some nodes to process significantly more data than others, creating performance bottlenecks.
Report an error in this question
Big Data AnalyticsHard
Q22. What is the difference between horizontal and vertical scaling?
- A.Vertical scaling adds additional machines to the cluster pool
- B.They are completely identical approaches to scaling infrastructure
- C.Horizontal scaling upgrades individual machine hardware resources
- D.Horizontal adds more machines; vertical adds more power to existing machines✓ Correct
Explanation
Horizontal scaling (scaling out) adds more machines to handle load. Vertical scaling (scaling up) adds more CPU, RAM, or storage to existing machines. Big data systems favor horizontal scaling.
Report an error in this question
Big Data AnalyticsHard
Q23. What is the concept of data lineage?
- A.The specific file format of the data such as CSV or JSON
- B.Tracking data from its origin through all transformations to its current state✓ Correct
- C.The chronological age of the data since it was created
- D.The total size of the data measured in bytes or rows
Explanation
Data lineage tracks the origin, movement, and transformation of data throughout its lifecycle, essential for governance, debugging, and compliance.
Report an error in this question
Big Data AnalyticsHard
Q24. What is the Lambda architecture?
- A.A single-server architecture running everything on one machine
- B.An architecture supporting only offline batch processing workloads
- C.An architecture supporting only real-time stream processing
- D.A big data architecture combining batch and real-time stream processing✓ Correct
Explanation
Lambda architecture processes big data using both a batch layer (for comprehensive, accurate views) and a speed layer (for real-time views), merging results in a serving layer.
Report an error in this question
Big Data AnalyticsHard
Q25. What is Apache Arrow?
- A.A cross-language platform for in-memory columnar data with zero-copy reads✓ Correct
- B.A distributed event streaming platform for real-time pipelines
- C.A supervised machine learning library for model training
- D.A standalone relational database for transactional query workloads
Explanation
Apache Arrow defines a language-independent columnar memory format for flat and hierarchical data, enabling efficient data interchange between systems with zero-copy reads.
Report an error in this question
Big Data AnalyticsHard
Q26. What is the Kappa architecture?
- A.A simplified Lambda alternative using only stream processing for real-time and batch✓ Correct
- B.An architecture designed only for small-scale single-machine data
- C.The exact same architecture as Lambda with no simplification
- D.An architecture supporting only offline batch processing workloads
Explanation
Kappa architecture simplifies Lambda by using a single stream processing layer for both real-time and reprocessed historical data, reducing complexity while maintaining the same capabilities.
Report an error in this question
Big Data AnalyticsHard
Q27. What is Apache Flink?
- A.A standalone relational database for transactional query workloads
- B.A stream processing framework with event-time processing and exactly-once semantics✓ Correct
- C.A batch-only data processor without any stream processing capability
- D.An interactive data visualization and dashboard charting tool
Explanation
Apache Flink is a distributed stream processing framework that handles both batch and stream processing with event-time semantics and exactly-once state consistency.
Report an error in this question
Big Data AnalyticsHard
Q28. What is eventual consistency in distributed systems?
- A.All replicas converge to the same value given enough time without new updates✓ Correct
- B.Consistency is fundamentally impossible in any distributed system
- C.All replicas are immediately and perfectly consistent at all times
- D.Data is never consistent across any of the distributed replicas
Explanation
Eventual consistency guarantees that if no new updates are made, all replicas will eventually return the same value, trading immediate consistency for availability and partition tolerance.
Report an error in this question
Big Data AnalyticsHard
Q29. What is the role of a distributed consensus algorithm like Raft or Paxos?
- A.Visualizing distributed system metrics on dashboards
- B.Sorting large distributed data records across nodes
- C.Training machine learning models on distributed nodes
- D.Ensuring all nodes agree on shared state despite failures✓ Correct
Explanation
Consensus algorithms like Raft and Paxos ensure that multiple nodes in a distributed system agree on a single value or state even when some nodes fail, critical for data consistency.
Report an error in this question
Big Data AnalyticsHard
Q30. What is columnar storage and why is it efficient for analytics?
- A.A type of secondary database index for accelerating query lookups
- B.Storing data by rows as in traditional relational database table layouts
- C.Storing data by columns not rows, enabling efficient compression and selective reads✓ Correct
- D.A data visualization technique for displaying column chart graphics
Explanation
Columnar storage stores each column separately, enabling better compression (similar values together) and efficient reads when queries only access specific columns, ideal for analytical workloads.
Report an error in this question