Research brief ·
Philippines payroll duplicate time-entry detection study
Research question: which evidence distinguishes a true duplicate time entry from two legitimate records that merely look alike?

Research finding
Count candidates and confirmed duplicates separately.
Methodology
This brief triangulates the headline measure against official Philippine government, regulatory, development, and labor sources. It translates the evidence into an operating control and separates context from recommendations.
| Measure | Interpretation |
|---|---|
| Candidate pairs are tested against source identity, work interval, approval lineage, correction history, and payroll treatment. | Context signal for planning; not a promise about an individual worker or provider. |
| 4 source records | Primary source links are listed and numbered below for review. |
Key takeaways
- Count candidates and confirmed duplicates separately.
- Review identity, interval, source, and correction lineage before changing a record.
- A detection rule proposes cases; an authorized owner decides treatment.
A duplicate is a conclusion, not a visual resemblance
Two rows with the same person, date, and hours may be duplicates.
They may also represent a split shift, two approved earning codes, a correction that preserves the original, or entries from separate source systems.
This study asks which evidence lets a Philippines-based payroll support team separate confirmed duplicates from legitimate lookalikes before payroll preparation.
The unit is a candidate pair or cluster.
The evidence scope includes stable worker identifier, employing entity, pay group, work date, start and end time where available, earning code, source-system identifier, creation time, correction reference, approval record, and intended payroll treatment.
The study does not determine whether hours are payable or whether an earning code is legally correct.
It examines whether the records support a duplicate classification.
Create the candidate set without silently editing it
Take a read-only extract from a defined period and record its source, extraction time, row count, and transformation steps.
Generate candidate pairs using several explicit rules.
Exact matches may compare worker, date, interval, and earning code.
Near matches may allow small differences in duration or timestamps.
Cross-system candidates may compare a timekeeping record with a manually entered adjustment.
Keep the rules separate so reviewers can see why each pair appeared.
Never delete, merge, or overwrite a source row during detection.
Assign a candidate identifier and retain both source references.
Sample at least three comparable cycles, including routine records, corrections, overnight work, split shifts, and multiple earning codes.
The resulting set measures what the rules flagged, not how many actual duplicates exist.
Adjudicate a sample with a written evidence standard
Before reviewing outcomes, define classifications such as confirmed duplicate, legitimate separate entry, superseded record, unresolved, and out of scope.
A confirmed duplicate should require evidence that two records describe the same intended work or payment event and that both would otherwise reach payroll treatment.
Similar values alone are insufficient.
Have two reviewers independently classify a sample that includes exact and near matches.
Record disagreement and route it to the authorized payroll owner.
Calculate agreement, confirmed-duplicate yield, false-positive share, and unresolved share for each detection rule.
If reviewers use information outside the defined evidence set, add that source to the method or mark the classification unsupported.
This makes the result reproducible and prevents intuition from changing from case to case.
Test population effects before recommending a rule
A rule that finds many candidates can still be poor if it floods the queue with legitimate entries.
Estimate how many records the rule reviews per thousand source rows and how much owner attention each confirmed case requires.
Examine false negatives through a separate risk-based sample, such as repeated identifiers, manual adjustments, overlapping intervals, or corrections entered close to cutoff.
No sample proves that all missed duplicates were found.
Report the selected population and uncertainty.
Also check whether a rule treats one entity, shift pattern, or worker group differently because its source data is structured differently.
The analysis should describe that data-quality difference rather than imply worker behavior.
Use findings to improve intake and review
Confirmed clusters should be traced to the earliest controllable point.
Repeated cross-system duplicates may indicate that a manual entry process lacks a source reference.
Superseded records that look active may indicate weak correction status.
Legitimate split shifts caught by an exact rule may need a distinguishing identifier.
Assign improvements to the source owner, integration owner, or payroll owner as appropriate.
A support specialist may run the detection, assemble candidate evidence, and record dispositions.
The specialist should not remove a time entry, change hours, or select pay treatment without approval.
Keep the candidate and decision records even after an authorized correction so later testing can determine whether the rule improved.
Document the rule so a second reviewer can repeat it
A detection rule needs more than a label such as exact match or possible duplicate.
Record the fields compared, normalization applied, tolerance, excluded values, source version, and run time.
State how blank identifiers, rounded times, overnight intervals, and reversed entries behave.
Include several de-identified examples that cover a confirmed duplicate, a legitimate match, and an unresolved pair.
Then ask a reviewer who did not build the rule to reproduce a small sample from the frozen extract.
Differences between the original and repeated result may reveal an undocumented filter or manual step.
Retain the rule version with each candidate set so later changes do not rewrite prior results.
When a rule changes, run old and new versions on the same bounded sample and report which candidates were added or lost.
A narrower rule may reduce workload but miss a class that owners still care about.
A broader rule may increase coverage while making the queue unusable.
The payroll owner should accept the operating tradeoff after reviewing both the detected cases and the missed-case sample.
This turns duplicate detection into a controlled, reviewable process instead of a spreadsheet formula whose assumptions disappear when its author leaves.
Limitations and evidence-led conclusion
Source systems use different identifiers, rounding conventions, overnight-shift logic, and correction models.
Extracts may omit details that a system interface displays.
A three-cycle study may not include seasonal schedules or rare adjustments.
Public labor and data-protection sources do not define an employer's timekeeping policy.
For those reasons, a candidate score is not proof of duplicate pay.
The evidence-led conclusion is to retain source identity, apply transparent candidate rules, classify a documented sample, and measure false positives alongside confirmed cases.
The best rule is not automatically the rule with the most alerts.
It is the rule that produces reviewable candidates at a workload the owner can inspect, while leaving ambiguous cases visible.
Outsourced support can organize that evidence and repeat the test.
Final changes to time or payroll records require the authorized owner.
Sources
FAQs
Does this research decide payroll treatment?
No. It studies operating evidence. An authorized payroll owner and qualified advisers must decide pay, tax, employment, privacy, and filing questions.
What may an outsourced payroll support specialist do?
The specialist may collect records, run documented comparisons, record exceptions, and prepare a review packet. The client-side owner retains interpretation, approval, and release authority.
For adjacent operating context, see Payroll Preparation and the payroll operations guide library.