How does cbcopy handle a duplicate value conflict when the DISTRIBUTED BY column is also the PRIMARY KEY? #2050
Replies: 1 comment
|
Hi @asheraz048i2c, I looked into this in the cbcopy source and tested it on a test cluster. Short answercbcopy does no pre-check, and it has no skip/log option for constraint violations. In a normal full copy, the duplicate rows are loaded and the primary key is silently not created. How cbcopy loads data:
What happens in a full copy (metadata + data):During the metadata phase, cbcopy runs DDL and data copy concurrently: each table is queued for data copy as soon as its CREATE TABLE runs. The ALTER TABLE ... ADD CONSTRAINT ... PRIMARY KEY statements run near the end of the pre-data restore. A failed metadata statement is logged as an error. I reproduced this with a table DISTRIBUTED BY (id) + PRIMARY KEY (id), 2,000,010 rows / 2,000,000 distinct ids: ERROR: ... ALTER TABLE ONLY public.t_dup ADD CONSTRAINT t_dup_pkey PRIMARY KEY (id); The result on the target:
|
Uh oh!
There was an error while loading. Please reload this page.
Context:
I'm migrating a production database from Greenplum 7 to Apache Cloudberry 2.0.1 using cbcopy. One of the tables (~880 million rows) has its DISTRIBUTED BY column also defined as the PRIMARY KEY.
Question:
If the source table contains duplicate values in this column — for example due to a constraint that was disabled or marked NOT VALID at some point on the source side — what is cbcopy's actual behavior when loading into a target table where the PRIMARY KEY constraint is active?
Specifically:
Does cbcopy do any pre-check on this column before starting the load, or does it just attempt the insert and let the target's constraint layer catch it?
When a duplicate key violation happens mid-load, does the entire table/partition load abort and roll back, or does it skip the conflicting row(s) and continue with the rest?
Is there any flag or option to control this behavior — for example, skipping and logging conflicting rows versus a hard failure?
I'd like a definitive answer before running this on a table this large in production, rather than relying on trial-and-error. Any insight from someone who has hit this scenario during a large-scale migration would help a lot.
Environment:
Source: Greenplum 7.x
Target: Apache Cloudberry 2.0.1
Table size: ~880 million rows
Column: DISTRIBUTED BY + PRIMARY KEY (same column)
All reactions