HomeSubjectsUniversityBlogAbout

Distributed Databases

Topic in Databases

210 total MCQsShowing 30 with explanations10 Easy10 Medium10 Hard

About This Topic

A distributed database is one logical database whose data is stored across several networked sites yet presented to users as a single system. Design questions cover horizontal, vertical and mixed fragmentation along with the completeness, reconstruction and disjointness rules, plus full versus partial replication. You should distinguish homogeneous from heterogeneous systems and know the forms of distribution transparency. For transactions, the two-phase commit protocol with its coordinator, prepare and commit phases, and blocking problem is essential, with three-phase commit as its extension. Query items focus on minimizing data transfer, often through semijoins, and theory items cover quorum-based replication and the CAP theorem.

Below are 30 practice questions from a pool of 210 Distributed Databases MCQs, one of 17 topics in Databases. Each shows the correct answer with an explanation; when you are ready, take a timed quiz to test recall under exam conditions.

Practice Questions

Each question below shows the correct answer with a full explanation. Use these to build conceptual understanding before attempting a timed quiz.

Distributed DatabasesEasy

Q1. A distributed database is:

  1. A.A database stored only in the cloud without any local presence
  2. B.A database stored on a single machine in one physical location
  3. C.A database spread across multiple sites connected by a network✓ Correct
  4. D.A backup copy of a database kept for disaster recovery purposes

Explanation

A distributed database has data distributed across multiple sites connected by a communication network.

Report an error in this question

Distributed DatabasesEasy

Q2. Data fragmentation in distributed databases means:

  1. A.Duplicating the entire database at every site simultaneously
  2. B.Compressing data to reduce storage requirements at each site
  3. C.Losing data during transfer between sites on the network link
  4. D.Dividing a relation into smaller parts stored at different sites✓ Correct

Explanation

Fragmentation divides a relation into fragments that are stored at different sites.

Report an error in this question

Distributed DatabasesEasy

Q3. Horizontal fragmentation divides a relation by:

  1. A.Rows and columns both
  2. B.Rows (tuples) only✓ Correct
  3. C.Neither rows nor cols
  4. D.Columns (attributes)

Explanation

Horizontal fragmentation splits a relation into subsets of tuples (rows).

Report an error in this question

Distributed DatabasesEasy

Q4. Vertical fragmentation divides a relation by:

  1. A.Rows (tuples) only
  2. B.Columns (attributes)✓ Correct
  3. C.Neither of the above
  4. D.Columns and rows both

Explanation

Vertical fragmentation splits a relation into subsets of attributes (columns), including the key in each fragment.

Report an error in this question

Distributed DatabasesEasy

Q5. Data replication in distributed databases means:

  1. A.Compressing data to reduce transfer size
  2. B.Deleting data from all distributed sites
  3. C.Encrypting data before sending it across
  4. D.Storing copies of data at multiple sites✓ Correct

Explanation

Replication stores copies of data fragments at multiple sites for availability and performance.

Report an error in this question

Distributed DatabasesEasy

Q6. The transparency goal in distributed databases means:

  1. A.Data cannot be distributed across multiple sites
  2. B.Users must know all data locations before querying
  3. C.Only administrators can query distributed data
  4. D.Users should not be aware that data is distributed✓ Correct

Explanation

Transparency hides the distribution details from users, making the system appear as a single database.

Report an error in this question

Distributed DatabasesEasy

Q7. Location transparency means:

  1. A.Data is always stored locally on the user's own machine
  2. B.Users must specify the exact data location in every query
  3. C.Users do not need to know where data is physically stored✓ Correct
  4. D.No remote access is possible from any site in the system

Explanation

Location transparency allows users to access data without knowing its physical location.

Report an error in this question

Distributed DatabasesEasy

Q8. A distributed DBMS (DDBMS) manages:

  1. A.Only local databases without any distribution
  2. B.A database distributed across multiple sites✓ Correct
  3. C.Only a centralized database on one machine
  4. D.Only cloud databases hosted by third parties

Explanation

A DDBMS manages the storage and retrieval of data across multiple networked sites.

Report an error in this question

Distributed DatabasesEasy

Q9. Which of the following is an advantage of distributed databases?

  1. A.Slower queries being preferred by users
  2. B.Higher costs being beneficial overall
  3. C.Improved reliability and availability✓ Correct
  4. D.Increased complexity as an advantage

Explanation

Distributed databases offer improved reliability (no single point of failure) and availability.

Report an error in this question

Distributed DatabasesEasy

Q10. Fragmentation transparency means:

  1. A.Users are unaware that data is fragmented✓ Correct
  2. B.Data cannot be fragmented at any site
  3. C.Only fragments can be queried directly
  4. D.Users must reassemble fragments manually

Explanation

Fragmentation transparency allows users to query data without knowing about its fragmentation.

Report an error in this question

Distributed DatabasesMedium

Q11. The two-phase commit (2PC) protocol ensures:

  1. A.Only read operations complete
  2. B.No atomicity guarantees at all
  3. C.Only local transaction atomicity
  4. D.Atomicity of distributed transactions✓ Correct

Explanation

2PC ensures that a distributed transaction either commits at all sites or aborts at all sites.

Report an error in this question

Distributed DatabasesMedium

Q12. In the 2PC protocol, the coordinator first sends:

  1. A.A prepare (vote) message to all participants✓ Correct
  2. B.A commit message directly to all the sites
  3. C.An abort message to cancel the transaction
  4. D.A query to all participants for data first

Explanation

In phase 1, the coordinator sends a prepare/vote-request message to all participant sites.

Report an error in this question

Distributed DatabasesMedium

Q13. If any participant votes No in 2PC, the coordinator:

  1. A.Ignores the vote and proceeds with commit
  2. B.Sends an abort message to all participants✓ Correct
  3. C.Retries the transaction automatically again
  4. D.Sends a commit message to all participants

Explanation

If any participant votes No (cannot commit), the coordinator sends a global abort to all participants.

Report an error in this question

Distributed DatabasesMedium

Q14. A distributed query involves:

  1. A.No data access from any site at all
  2. B.Only local data access on one site
  3. C.Accessing data from a single table
  4. D.Accessing data from multiple sites✓ Correct

Explanation

A distributed query requires accessing and combining data from tables stored at different sites.

Report an error in this question

Distributed DatabasesMedium

Q15. The CAP theorem states that a distributed system can guarantee at most:

  1. A.Two of three: Consistency, Availability, Partition Tolerance✓ Correct
  2. B.None of these properties can be guaranteed in any system
  3. C.Only one property out of the three possible in the theorem
  4. D.All three properties at the same time without any trade-offs

Explanation

The CAP theorem (Brewer's theorem) states that it's impossible to simultaneously guarantee all three properties.

Report an error in this question

Distributed DatabasesMedium

Q16. Semi-join optimization in distributed queries:

  1. A.Increases data transfer by sending all records across the link
  2. B.Uses full Cartesian products between all fragments in the join
  3. C.Reduces data transfer by sending only relevant join attributes✓ Correct
  4. D.Eliminates all joins from the distributed query execution plan

Explanation

Semi-join reduces network data transfer by first sending only the join attributes to filter matching tuples.

Report an error in this question

Distributed DatabasesMedium

Q17. Mixed fragmentation combines:

  1. A.No fragmentation is used at all
  2. B.Only vertical fragmentation types
  3. C.Horizontal and vertical fragmentation✓ Correct
  4. D.Only horizontal fragmentation types

Explanation

Mixed (hybrid) fragmentation applies both horizontal and vertical fragmentation strategies.

Report an error in this question

Distributed DatabasesMedium

Q18. Full replication means:

  1. A.Every site has a complete copy of the entire database✓ Correct
  2. B.No replication exists at any site in the system at all
  3. C.Only one site has the data and others have no copies
  4. D.Data is partially replicated at only a few of sites

Explanation

Full replication stores a complete copy of the database at every site.

Report an error in this question

Distributed DatabasesMedium

Q19. The blocking problem in 2PC occurs when:

  1. A.No failures occur during the execution of the two-phase commit protocol
  2. B.The coordinator fails after sending prepare but before sending the decision✓ Correct
  3. C.All participants commit successfully without any failures in the protocol
  4. D.The network is fast and reliable with no message delays or lost packets

Explanation

Blocking occurs when the coordinator crashes after prepare, leaving participants uncertain about whether to commit or abort.

Report an error in this question

Distributed DatabasesMedium

Q20. Distributed deadlock detection is more complex because:

  1. A.All transactions are local and never span multiple sites
  2. B.Deadlocks can involve transactions at multiple sites✓ Correct
  3. C.Deadlocks never occur in distributed database systems
  4. D.Only one site has transactions running at any time

Explanation

Distributed deadlocks span multiple sites, requiring global wait-for graph construction or timeout-based detection.

Report an error in this question

Distributed DatabasesHard

Q21. The three-phase commit (3PC) protocol improves on 2PC by:

  1. A.Using only one phase for the commit decision
  2. B.Removing the prepare phase from the protocol
  3. C.Adding a pre-commit phase to avoid blocking✓ Correct
  4. D.Ignoring all failures that occur at any site

Explanation

3PC adds a pre-commit phase between prepare and commit, avoiding the blocking problem of 2PC under certain failures.

Report an error in this question

Distributed DatabasesHard

Q22. In distributed databases, the global query optimization must consider:

  1. A.Only CPU time without considering disk or network access costs
  2. B.Only disk I/O costs without considering network or CPU overhead
  3. C.Only local processing costs without considering network overhead
  4. D.Network communication costs in addition to local processing costs✓ Correct

Explanation

Distributed query optimization must account for network transfer costs, which can dominate total cost.

Report an error in this question

Distributed DatabasesHard

Q23. Paxos is a protocol used for:

  1. A.Achieving consensus in distributed systems✓ Correct
  2. B.Indexing data for faster retrieval queries
  3. C.Query optimization in database management
  4. D.Schema design for database normalization

Explanation

Paxos is a consensus protocol that enables distributed nodes to agree on a value despite failures.

Report an error in this question

Distributed DatabasesHard

Q24. The RAFT consensus algorithm is designed to be:

  1. A.Faster than all other protocols in every distributed environment
  2. B.A replacement for SQL in distributed relational database systems
  3. C.More understandable than Paxos while providing the same guarantees✓ Correct
  4. D.A type of index for organizing data in distributed hash tables

Explanation

RAFT was designed as a more understandable alternative to Paxos for distributed consensus.

Report an error in this question

Distributed DatabasesHard

Q25. In a federated database system:

  1. A.No integration is possible between any databases at all
  2. B.Autonomous databases cooperate to provide integrated access✓ Correct
  3. C.Only one database exists in the entire distributed system
  4. D.All databases are identical copies of the same data schema

Explanation

A federated database system integrates multiple autonomous databases while preserving their independence.

Report an error in this question

Distributed DatabasesHard

Q26. The eventual consistency model guarantees that:

  1. A.No consistency is provided between any of the replicas in the system
  2. B.Immediate consistency at all times for every read at every site node
  3. C.All reads always return the latest written value at every site instantly
  4. D.All replicas will converge to the same value if no new updates are made✓ Correct

Explanation

Eventual consistency guarantees that replicas will converge to the same value eventually if updates stop.

Report an error in this question

Distributed DatabasesHard

Q27. The quorum-based protocol requires:

  1. A.Qr = Qw = 1 for consistency meaning only one replica needs to respond
  2. B.Read quorum (Qr) + Write quorum (Qw) > total replicas (N) for consistency✓ Correct
  3. C.No quorum is needed and any single replica can respond to operations
  4. D.Qr + Qw less than N for consistency between read and write quorums

Explanation

Quorum protocols require Qr + Qw > N and 2*Qw > N to ensure read-write and write-write consistency.

Report an error in this question

Distributed DatabasesHard

Q28. Vector clocks in distributed systems are used to:

  1. A.Capture causal relationships between events at different sites✓ Correct
  2. B.Measure CPU speed and processing capacity at each network node
  3. C.Optimize queries by selecting the best execution plan for joins
  4. D.Tell wall-clock time accurately across all distributed node sites

Explanation

Vector clocks track causality between events in a distributed system, determining potential concurrency.

Report an error in this question

Distributed DatabasesHard

Q29. Distributed hash tables (DHTs) are used in:

  1. A.Peer-to-peer distributed storage systems✓ Correct
  2. B.Centralized databases on a single machine
  3. C.Single-machine databases without networks
  4. D.Only for indexing within one database node

Explanation

DHTs distribute key-value storage across nodes in a peer-to-peer network for scalable lookups.

Report an error in this question

Distributed DatabasesHard

Q30. The coordinator selection problem in distributed databases can be solved by:

  1. A.Manual intervention always requiring a human administrator
  2. B.No algorithm exists for solving the coordinator selection
  3. C.Random selection only without any formal algorithmic approach
  4. D.Election algorithms like the Bully algorithm or Ring algorithm✓ Correct

Explanation

Election algorithms (Bully, Ring) automatically select a new coordinator when the current one fails.

Report an error in this question

Ready to test yourself on Distributed Databases?

Take a timed quiz drawn from 210+ questions on this topic. No signup required — your progress saves in your browser.

Start Distributed Databases Quiz