Compare translations from different model configurations. Choose the better translation for each sample.
Show differences between translations like GitHub
Choose 'Can't Tell' if both translations are equally good/bad or you can't decide