CrowdLogging: Distributed, Private, and Anonymous Search Logging. Feild, H., Allan, J., & Glatt, J. In Proceedings of the 34th international ACM SIGIR Conference on Research and development in Information Retrieval (SIGIR 2011), pages 375–384, 2011. ACM. Issue: 2
doi  abstract   bibtex   
We describe CrowdLogging, an approach for distributed search log collection, storage, and mining, with the dual goals of preserving privacy and making the mined informa- tion broadly available. Most search log mining approaches and most privacy enhancing schemes have focused on cen- tralized search logs and methods for disseminating them to third parties. In our approach, a user’s search log is en- crypted and shared in such a way that (a) the source of a search behavior artifact, such as a query, is unknown and (b) extremely rare artifacts—that is, artifacts more likely to contain private information—are not revealed. The ap- proach works with any search behavior artifact that can be extracted from a search log, including queries, query reformulations, and query-click pairs. In this work, we: (1) present a distributed search log collection, storage, and mining framework; (2) compare several privacy policies, including differential privacy, showing the trade-offs between strong guarantees and the utility of the released data; (3) demonstrate the impact of our approach using two existing research query logs; and (4) describe a pilot study for which we implemented a version of the framework.
@inproceedings{feild_crowdlogging_2011,
	title = {{CrowdLogging}: {Distributed}, {Private}, and {Anonymous} {Search} {Logging}},
	isbn = {978-1-4503-0757-4},
	doi = {http://doi.acm.org/10.1145/2009916.2009969},
	abstract = {We describe CrowdLogging, an approach for distributed search log collection, storage, and mining, with the dual goals of preserving privacy and making the mined informa- tion broadly available. Most search log mining approaches and most privacy enhancing schemes have focused on cen- tralized search logs and methods for disseminating them to third parties. In our approach, a user’s search log is en- crypted and shared in such a way that (a) the source of a search behavior artifact, such as a query, is unknown and (b) extremely rare artifacts—that is, artifacts more likely to contain private information—are not revealed. The ap- proach works with any search behavior artifact that can be extracted from a search log, including queries, query reformulations, and query-click pairs. In this work, we: (1) present a distributed search log collection, storage, and mining framework; (2) compare several privacy policies, including differential privacy, showing the trade-offs between strong guarantees and the utility of the released data; (3) demonstrate the impact of our approach using two existing research query logs; and (4) describe a pilot study for which we implemented a version of the framework.},
	booktitle = {Proceedings of the 34th international {ACM} {SIGIR} {Conference} on {Research} and development in {Information} {Retrieval} ({SIGIR} 2011)},
	publisher = {ACM},
	author = {Feild, Henry and Allan, James and Glatt, Joshua},
	year = {2011},
	note = {Issue: 2},
	keywords = {CrowdLogging, Distributed query logs, crowdlogging, log analysis, logfiles, private query log analysis, query logs, web logs: query logs},
	pages = {375--384},
}

Downloads: 0