Ally Zhang scite author profile

We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges. The data contains 107,292 examples of English sentences paired with web photographs. The task is to determine whether a natural language caption is true about a pair of photographs. We crowdsource the data using sets of visually rich images and a compare-and-contrast task to elicit linguistically diverse language. Qualitative analysis shows the data requires compositional joint reasoning, including about quantities, comparisons, and relations. Evaluation using state-of-the-art visual reasoning methods shows the data presents a strong challenge. * Contributed equally. † Work done as an undergraduate at Cornell University. 1 In parts of this paper, we use the term compositional differently than it is commonly used in linguistics to refer to reasoning that requires composition. This type of reasoning often manifests itself in highly compositional language.The left image contains twice the number of dogs as the right image, and at least two dogs in total are standing.One image shows exactly two brown acorns in back-to-back caps on green foliage.

show abstract

Evaluating Models’ Local Decision Boundaries via Contrast Sets

Gardner

Artzi

Basmov³

et al. 2020

194

182

View full text Add to dashboard Cite

Standard test sets for supervised learning evaluate in-distribution generalization. Unfortunately, when a dataset has systematic gaps (e.g., annotation artifacts), these evaluations are misleading: a model can learn simple decision rules that perform well on the test set but do not capture the abilities a dataset is intended to test. We propose a more rigorous annotation paradigm for NLP that helps to close systematic gaps in the test data. In particular, after a dataset is constructed, we recommend that the dataset authors manually perturb the test instances in small but meaningful ways that (typically) change the gold label, creating contrast sets. Contrast sets provide a local view of a model's decision boundary, which can be used to more accurately evaluate a model's true linguistic capabilities. We demonstrate the efficacy of contrast sets by creating them for 10 diverse NLP datasets (e.g., DROP reading comprehension, UD parsing, and IMDb sentiment analysis). Although our contrast sets are not explicitly adversarial, model performance is significantly lower on them than on the original test sets-up to 25% in some cases. We release our contrast sets as new evaluation benchmarks and encourage future dataset construction efforts to follow similar annotation processes.

show abstract

A Corpus for Reasoning About Natural Language Grounded in Photographs

Suhr¹,

Zhou²,

Zhang³

et al. 2018

Preprint

View full text Add to dashboard Cite

Evaluating Models' Local Decision Boundaries via Contrast Sets

Gardner

Artzi²,

Basmova³

et al. 2020

Preprint

View full text Add to dashboard Cite

Best Friend or Worst Enemy? -- Dynamics and Multiple Equilibria with Arbitrage, Production and Collateral Constraints

Zhang

2017

SSRN Journal

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

customersupport@researchsolutions.com

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Ally Zhang

A Corpus for Reasoning about Natural Language Grounded in Photographs

Evaluating Models’ Local Decision Boundaries via Contrast Sets

A Corpus for Reasoning About Natural Language Grounded in Photographs

Evaluating Models' Local Decision Boundaries via Contrast Sets

Best Friend or Worst Enemy? -- Dynamics and Multiple Equilibria with Arbitrage, Production and Collateral Constraints

Contact Info

Product

Resources

About